Skip to main content
Compliance

A Frontier Model Beat Our Extractor 97.9% to 68.2%. We Shipped Ours Anyway.

Sixty clinical documents, 516 gold-labeled items, three trials. The frontier model out-recalled our deterministic extractor 97.9% to 68.2% — and still could not ship. Here is the whole benchmark, including the number that goes against us.

The benchmark nobody publishes

Vendors publish the benchmark they win. We ran one we lost, and the losing number is on the first page of the document it belongs to. Here is the whole thing, because the shape of the result is more useful than the score.

The task: take a hospital discharge summary and pull out the things a family actually has to act on — which medicines started, stopped or changed; which appointments exist; which warning signs mean call someone; what the patient may not do yet. We built a deterministic extractor for it: rule-governed clause matching over a curated, cited knowledge base, no generation anywhere in the path.

The corpus: 60 discharge notes carrying 516 hand-labeled gold items. Twenty of those notes were written and gold-labeled before the engine ever ran against them, and the engine was not touched afterwards.

The result

The frontier model won. On identical notes, the same four categories, the same matcher, three trials: 97.9% recall against our 68.2%. On medication changes specifically — the category that matters most in the first week home — ours managed 56.6%.

It was not a narrow loss and we are not going to describe it as one.

It was also more careful than we expected

The easy story would be that the model hallucinated and we caught it. That is not what happened.

  • Every source span it cited was actually present in the note. No invented quotes.
  • It preserved 39 of 39 numeric thresholds on every trial — the “gain more than 3 pounds in a day” kind of number, the kind that kills people when it drifts.
  • Across 180 note-runs it introduced exactly one number the discharge did not contain: 911.

If your mental model of a 2026 frontier model is a confident fabricator, update it. On this task it was disciplined.

So why did we ship the other one?

Two reasons, and only one of them is about accuracy.

1. It was not the same tool twice

Run identical input through it three times and it returned a different item set on 11 of 60 notes. Not wrong — different. For most software that is a curiosity. For a page a family reads at 2am, and for anything that has to be validated, reproduced, or audited, it is disqualifying. You cannot certify a function that does not agree with itself.

A system that gives a different answer to the same question is not a system with a recall problem. It is a system without a defined output.

2. It out-recalled partly by saying more

Against 516 gold items it printed 657; our engine printed 350. On warning signs it emitted 232 items against our 106. Higher recall, yes — and a 6.1% wrong-item rate against our 2.6%. Some of that lift is genuine retrieval. Some of it is volume. If you only report recall, you cannot tell those apart, which is why we report recall and false-extraction together, always, per category.

What this actually means, outside of clinical text

The finding generalizes to anything where the output is consumed as an instruction rather than as a draft:

  • Determinism is a feature, and it is one generation does not currently offer. If your output feeds an audit, an accreditation, a regulated filing, or a safety instruction, run the same input twice before you believe any benchmark.
  • Recall alone is a marketing number. Report the false-positive rate beside it or the comparison is meaningless.
  • Verbatim reproduction is a design choice with a measurable cost. Ours lifts a warning sentence exactly as the clinician wrote it. That caps our recall — and it means a threshold cannot be softened in paraphrase.
  • Blind splits, or nothing. Twenty of our notes were labeled before the engine saw them. Without that, a benchmark measures how well you tuned, not how well you generalize.

Where the frontier model is the right answer

We are not arguing against generation. On the same evidence we would use it for drafting where a human reviews every line, for coverage sweeps where over-production is cheap and misses are expensive, and for anything exploratory. Its recall advantage here was real and large.

What we would not do is put it on the last mile of a safety instruction with no human between it and the reader — not because it is careless, but because it is not reproducible, and reproducibility is what the last mile requires.

One more number, because it cuts against us too

All 60 of our FHIR R4 exports validate with zero errors against a public server we do not control. That server first found four defects in our own export. We would not have found them ourselves, and we would not have known to look. External validation is not a formality; it is the only part of your test suite you did not write.

Method, so you can argue with it

60 discharge notes, 516 gold-labeled items, four categories. Twenty notes written and labeled blind before the engine ran; engine unchanged after. Frontier model given the same four buckets and scored by the identical matcher, three trials. Recall and false-extraction reported together, per category. Machine-written logs retained for every run.

The point

We publish the benchmark we lost because the number that goes against you is the only one a reader cannot get anywhere else — and because a firm that hides a 68.2% is a firm that will hide something worse later. If you are evaluating an AI system for federal or clinical deployment, ask the vendor for the measurement where their approach came second. If they do not have one, they did not run it.

1 business day response

Building an agentic compliance capability?

We design and deploy agentic compliance tooling that stays inside a FedRAMP boundary and actually ships measurable time savings.

Talk to usRead more insights →
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE