CustomsGenius

We check our duty math against CBP's final answer

CustomsGenius validates its duty recomputation against CBP's own liquidation outcomes, and publishes the method, the numbers, and the limitations.

Numbers regenerated at publication · 2026-07-18

Every refund estimate on our platform is, at bottom, a duty recomputation: what an entry should have been assessed under the tariff schedule and every applicable program, versus what was paid. For entries CBP has finally liquidated, there is a ground truth for that arithmetic: the final duty total CBP itself posted when it closed the entry. We run that comparison continuously, entry by entry, and we publish the method, the numbers, and the limitations.

To our knowledge (as of July 2026, re-checked at publication), no other tariff analytics product publishes a validation of its duty computation against ACE liquidation outcomes. Accuracy claims are common in this industry; published, reproducible validation methods are not. We publish ours (including the parts that are not yet flattering) because a refund claim is only as strong as the arithmetic behind it.

We do not claim perfect accuracy, and you should be skeptical of anyone who does. We claim a measured, reproducible agreement record, stated below with its exact scope.

Why this matters for a refund claim

A refund vehicle (a CAPE declaration, a post-summary correction, a protest) asks CBP to agree that the duty on an entry should have been X, not the Y that was paid. The number X comes out of a computation engine.

Most tariff software is checked, at best, against published rate tables. But assessment behavior lives in the details those tables don't settle by themselves: effective-date boundaries, Chapter 99 routing, program stacking, country scope. Liquidation is CBP's own final computation on a real entry: the same authority that will adjudicate your claim, doing the same arithmetic your claim depends on. It is the strictest oracle available, so it is the one we test against.

The method, in plain language

Our customers upload their own ACE entry summary data. For every entry in that corpus that CBP has actually liquidated, we compare three entry-level duty totals:

  1. Liquidated: CBP's final total duty from the ACE entry summary;
  2. Filed: the duty declared at entry, summed across the entry's lines at program grain (MFN / Section 232 / Section 301 / Section 122 / IEEPA);
  3. Engine: a rates-only re-simulation of each line (HTS-10, origin, entry date, SPI) through the CustomsGenius tariff engine.

Two totals agree when they differ by at most $1 absolute or 0.5% relative, whichever is larger. Those tolerances are constants imported from the engine codebase (the same values the product uses), not thresholds tuned per report.

Every comparable entry is then classified into exactly one agreement class:

ClassMeaning
LIQ_CONFIRMS_FILEDCBP liquidated at the filed amount and the engine agrees (liq ≈ filed ≈ engine)
LIQ_CONFIRMS_ENGINECBP changed the duty at liquidation and the final amount matches the engine, not the filing
LIQ_CONFIRMS_FILED_ENGINE_DIFFERSCBP liquidated as filed but the engine expects a different total
ALL_THREE_DIFFERno two of the three totals agree within tolerance
LIQ_ZERO_OR_SUSPENDED$0 or non-final liquidation: set aside, never counted as agreement

Entries that cannot be compared honestly are excluded before classification: no liquidation data, not yet liquidated, or "incomparable" (duty outside the covered program grain, a non-ad-valorem MFN rate, or the partial-knowledge case below). Exclusions are never counted as agreement.

Results are persisted at cluster grain (HTS × month × class) with no customer entry numbers: the validation ledger carries no customer-identifying data.

The precise rules we hold ourselves to

These are the rules exactly as implemented; they are what keep the numbers below honest.

Current numbers

All tables on this page are machine-generated from the production validation ledger and dated. Nothing here is hand-computed.

Three-way liquidation comparison · regenerated 2026-07-18 · corpus: 7 entry months (2019-06, 2020-09, 2020-11, 2021-06, 2022-06, 2022-11, 2023-06), replay runs through 2026-07-17

WhatNumber
Entries classified (final liquidation on file, comparable)14,659
CBP's final liquidated total ≈ the filed total (filed axis)14,417 / 14,659 (98.3%)
CBP's final liquidated total ≈ our engine total (engine axis)8,519 / 14,659 (58.1%)
Entries where CBP changed the duty at liquidation and the final amount matched our engine, not the filing149
Entries where no two of the three totals agree93
Set aside as $0-liquidation or non-final (never counted as agreement)7,501

Read those two axes correctly. The filed axis says CBP overwhelmingly liquidated these classified entries at the filed amount: a statement about CBP's outcomes versus the filings, and about the integrity of the comparison pipeline; it is not a claim about our engine. The engine axis is the claim about our engine: on 8,519 of 14,659 classified entries, CBP's final liquidated total lands within tolerance of our independent recomputation, including 149 entries where CBP corrected the duty at liquidation and the corrected amount matched our engine rather than the original filing.

The remaining engine-axis gap is dominated by the "CBP liquidated as filed but the engine differs" class (6,047 entries): the boundary-precision and exclusion-take-up limits described above, plus the Section 232 content-basis uncertainty listed under limitations. Under those rules a large share of that residual is a limit of what entry data can express, not a measured engine error, but we cannot cleanly separate the two from this data alone, so we publish the whole classified universe, per month, rather than a curated subset:

Per entry month · same generation (2026-07-18) · counts derived from the per-class table of the same run

Entry monthClassifiedliq ≈ filedliq ≈ engineEngine axis
2019-0688450.0%
2020-091212541.7%
2020-112,1962,1961,00345.7%
2021-062,5432,5431,60163.0%
2022-062,8422,8421,67859.0%
2022-113,4623,4622,11461.1%
2023-063,5963,3542,11458.8%
All months14,65914,4178,51958.1%

2019-06 and 2020-09 are small months (8 and 12 classified entries); they are shown because the rule here is the whole classified universe, not the months that look best.

The wider evidence ledger

The three-way comparison is one oracle among several we run against the same engine. The evidence ledger sweeps the engine's whole answer surface and classifies every cell by the strongest evidence touching it:

Evidence ledger · generated 2026-07-17

EvidenceNumber
Single-HTS liquidation clusters where CBP's final liquidated money confirms the engine (entry months 2019-06 → 2023-06)894
Entry-replay agreements: declared duty on real filed entries matched the engine within tolerance (8 entry months, 2019-06 → 2026-06)5,007
Entries where CBP's own CAPE refund recompute matched the engine's IEEPA recompute (of 1,385 compared; 590 attributable to a single HTS/origin under our attribution rules)1,130
Share of the engine's current answer surface (2,425,764 HTS × country cells) backed by evidence beyond the engine itself: an oracle point, primary-authority citation, partial authority coverage, or a tracked open dispute91.17%
The same share for the skeptical reader, excluding tracked open disputes73.02%
Open disagreements we enumerate rather than hide (authority findings / replay clusters / mirror diffs)313 / 758 / 13

That 91.17% is an evidence-coverage number, not an accuracy number: it measures how much of the answer surface is supported by something other than our own code, never how often we are right. The engine's answer surface is swept exhaustively (23,782 HTS-10 codes × 102 country equivalence classes × 1,127 effective-date boundaries back to 2018), and each cell is classified by the strongest evidence touching it, with open disputes counted and listed, not netted away. Quote it only next to its companion: excluding tracked open disputes, the share is 73.02%.

What this does, and does not, claim

It claims exactly this: within the classified universe described above, CBP's final liquidated entry totals agree or disagree with the filed and engine totals per the class table, within the stated tolerance. Where the liquidated total matches the engine total, the engine's entry-total duty computation is corroborated by CBP's own final assessment.

It does not claim:

Honest limitations

How to audit us

The methodology is reproducible, and that is deliberate.

Demand the same treatment for your own data

Give us your ACE entry summary data and we will run the identical three-way comparison (your filings, CBP's liquidated totals, our recompute) and hand you the per-entry classification under exactly the rules on this page.

Book a walkthrough

If you find a liquidated entry our engine gets wrong, we want it: that is one more oracle point, whichever way it cuts.

Numbers on this page are regenerated from production validation data at publication and dated in place. Last regeneration: 2026-07-18.

Request Beta Access

Get early access to CustomsGenius and start recovering IEEPA refunds faster.

Beta Pilot Ongoing