Talk about a build

AI and evidence 2026-06-22 4 min

Why we let a human overrule the adversarial checker

We measured what would happen if our skeptical second pass ran automatically. It would have demoted the strongest competitor most and handed our client back first place. A checker tuned to argue, applied mechanically, recreates the bias it was built to remove.

Scott Hodson

We built an adversarial second pass to catch the fabrication classes our quote verification cannot see. It reads every claim against its own evidence and tries to refute it.

The obvious next step is to wire it up: let its objections automatically demote claims and adjust scores. Full automation, no human in the loop, no opportunity for someone to wave through an inconvenient objection.

We measured what that would do before shipping it. It would have made the output worse in a specific and instructive way.

What the measurement showed

Applied mechanically, the adversarial pass would have demoted the strongest competitor far more than it demoted the client, and handed the client back the top spot.

The client had just been moved out of first place by a parity fix. The automation would have quietly undone that.

Why it happened

Not because the checker was wrong. On any individual objection it was usually reasonable.

The mechanism is arithmetic.

The strongest competitor had the most evidence, because they were the most active company in the market. More content, more campaigns, more public statements, more surface area. More evidence produced more claims. More claims gave a checker tuned to find fault more things to find fault with.

Objection count scaled with evidence volume. Evidence volume scaled with how much the company was actually doing. So the companies doing the most accumulated the most objections, and mechanical demotion converted activity into penalty.

The checker was measuring how much there was to argue with. Applied automatically, that becomes a score.

The general shape of the failure

This generalises well beyond our case, and it is worth naming clearly:

A checker tuned to argue, applied mechanically, reintroduces the exact bias it was built to remove.

The adversarial pass exists to counteract a bias toward flattering conclusions. Automate its output and you get a new bias, in a different direction, with the same structural cause: a systematic relationship between the volume of material and the volume of criticism.

Any review process with an asymmetric objective has this property. A checker whose job is to find problems finds more problems where there is more to look at. That is not a defect in the checker. It is what “tuned to find problems” means.

The error is treating its output as a measurement rather than as a prompt.

What we do instead

Warn, do not block.

The adversarial pass raises objections. A human reads them and decides. Where a human overrules an objection, the objection is still published beside the claim, with the reason it was overruled.

Three properties follow.

The reader sees the disagreement. They are not asked to trust that it was resolved correctly. They can read the claim, read the objection, read the reasoning, and form their own view. That is more information than a report that only shows resolved conclusions.

The human is accountable in public. Overruling an objection means writing down why, in a document the client reads. That is a meaningful constraint on doing it carelessly, and a considerably better one than an approval button.

The bias does not compound. An objection that would have mechanically moved a score instead becomes a note. Notes do not accumulate into a systematic distortion the way automatic adjustments do.

The cost, stated honestly

This is slower. It requires a person with judgement to read objections and make calls, and that person is a bottleneck.

It also means the pipeline is not fully automated, which is a slightly awkward thing to admit in a product that emphasises automation elsewhere. Data collection is automated. Scanning is automated. Scoring against published requirements is automated. This one step is not, and we would rather say so than imply the whole thing runs untouched.

The alternative was measurably worse output. That is the trade, and we would rather have a slower system that is right than a fully automatic one that has a known bias with a known direction.

The broader lesson for AI systems

There is a strong pull toward closing the loop. If you have built a checker, wiring its output directly into the thing it checks feels like the finished version. Anything less feels like an unfinished automation.

That instinct is worth resisting until you have measured what the closed loop actually does.

Two questions are worth asking of any automated critic before you let it act:

Does its output correlate with something other than the thing it is supposed to measure? Ours correlated with evidence volume. Yours might correlate with document length, verbosity, domain familiarity, or how well-represented a subject is in training data. Any of those becomes a bias the moment the output is applied automatically.

What does it do to the cases at the extremes? Our failure was invisible in the middle of the distribution and severe at the top. Testing on typical cases would have shown nothing wrong.

The useful move is to run the automation in shadow mode, compare its decisions against the human ones, and look specifically at where they diverge most. That divergence is where the structural bias lives.

We did that, found the bias, and shipped the version with a human in it.

The full description of all four controls, and what each of them is there to prevent, is on how we prove it.

Find out what your marketing
is actually achieving.

One conversation, no obligation: what you sell, who you lose deals to, and what you cannot currently see. If a build is a fit, it is scoped and quoted from there. There is no signup and no self-serve tier.

No prices on this site. Every engagement is scoped and quoted after the first meeting.