Four error modes, two of which fight each other
When a coding manager is asked about accuracy, the answer is usually a single percentage from an internal audit. That percentage collapses four different failures into one figure, and because two of the four have opposite financial and regulatory consequences, the figure can stay flat while the situation gets substantially better or substantially worse. Any serious conversation about coding accuracy starts by separating them.

The wrong code. A code was assigned that does not describe what happened. Usually a specificity error or a near-neighbor in the same family. It may pay or it may deny, and it corrupts every downstream analysis of what the organization actually does.
The missing code. Something documented and codeable was not coded. Lost payment, and in risk-adjusted arrangements, lost condition capture that also distorts the clinical picture carried forward.
The unsupported code. A code was assigned that the documentation does not support. This is the one with regulatory consequences: payer audit programs and government auditors review after payment, and repayment obligations attach. Its cost is not the payment, it is the exposure.
Structural errors. Sequencing, modifiers, bundling, laterality, the parts that are mechanical and rule-driven rather than interpretive. Cheapest to detect, most amenable to deterministic edits, and disproportionately represented in denials.
The second and third are the pair that fight. Tightening rules to eliminate unsupported codes pushes coders toward conservatism, which increases missed codes. Loosening to capture more increases exposure. There is no setting that minimizes both, and an organization that has not decided where it wants to sit on that trade-off will oscillate — tighten after an audit, loosen after a revenue review, and never converge.
You are probably here because
- Volume grew, the coding queue grew with it, and headcount is not available
- A vendor is quoting an accuracy figure and you cannot tell what it was measured against
- An audit found unsupported codes and the corrective action slowed everything down
- You are building a coding product and need an evaluation a customer will believe
The measurement section is the core of this. The vendor-evaluation section is what to send to a salesperson. If your problem is capacity rather than accuracy, the section on where coder time actually goes is the one that will help.
Most coding errors are documentation errors
A coder can only code what is written. If the note does not state laterality, does not link the finding to a diagnosis, does not record the elements the guideline requires, or records them in a place the coder does not see, no amount of coder skill produces the right code. It produces a query, which produces a delay, which produces the queue everyone is trying to solve with headcount.
This has a practical consequence that reorders most improvement projects: the highest-yield analysis is not an audit of coders, it is an analysis of queries. Group physician queries by the documentation element that was missing, by template, by specialty and by author. The result is nearly always a short list of specific, fixable defects — a note template that does not prompt for something, a specialty where the guideline changed and nobody rebuilt the template, a handful of clinicians with a habit that generates a query on most of their charts.
Fixing the top few query causes reduces coder rework, shortens the cycle, and improves accuracy at the same time, without a trade-off between the second and third error modes. It is the rare intervention in this domain with no downside, and it is usually available because the query data exists and nobody has grouped it.
| Error mode | What it actually costs | Where it originates |
|---|---|---|
| Wrong code | Denials, rework, and corrupted internal analytics | Ambiguous documentation, or a guideline change nobody trained on |
| Missing code | Payment not collected; in risk arrangements, condition capture lost | Documentation that never stated it, or a conservative posture after an audit |
| Unsupported code | Repayment exposure, and audit scrutiny that outlasts the finding | Templates that assert more than was done; pressure on productivity targets |
| Structural error | Denials, high volume, low dollar value each | Missing edits before submission. Mechanical, and mechanically preventable |
You cannot measure accuracy against one auditor
Here is the fact that most accuracy discussions skip. Experienced, credentialed coders do not fully agree with each other on complex charts. Agreement is high on straightforward encounters and falls as complexity rises, particularly where the code depends on interpreting clinical intent from prose. This is not a knock on coders; it is a property of a task where the input is natural language written for another purpose.
It has two consequences, and both are frequently ignored.
First, an accuracy figure measured against a single auditor is measured partly against that auditor's idiosyncrasies. If you want a number that means something, the reference has to be a set adjudicated by more than one qualified coder, with disagreements resolved explicitly and the disagreement rate itself recorded.
Second, that disagreement rate is the practical ceiling. If two qualified coders differ on a meaningful share of your complex charts, no system — human or otherwise — can be scored much above that against a single-reviewer standard, and any vendor quoting a figure far above it is telling you about the composition of their evaluation set rather than about your charts.
A usable evaluation set has five properties. It is drawn from your charts, in your specialty mix. It is stratified by complexity, so straightforward encounters do not drown the hard cases that determine whether a system is usable. It is double-coded and adjudicated, with the disagreements kept rather than discarded. It reports per code family, because a system can be excellent on one and unusable on another. And it reports direction — over versus under — separately, because the two are not equivalent and averaging them hides the one that carries regulatory exposure.
Where automation genuinely works today
Automated coding is not one capability. It works well in some domains and poorly in others, and the difference is predictable from three properties: how structured the source document is, how narrow the code space is, and whether the code depends on anything not written down.
How ready each domain is for meaningful automation — our read
Our engineering judgment of what holds up unattended in production, not a benchmark. Every row is worth testing on your own charts before believing it.
The top rows deserve attention because they are undersold. Structural edits — modifier logic, bundling checks, laterality consistency, sequencing rules — are deterministic, auditable and cheap, and they eliminate a category of denial that is high in volume. Many organizations run some of these inside a billing system and never look at which ones are enabled or how often each fires.
Query detection is the quietly valuable one. A system that reads a note before it reaches a coder and flags the missing element that will generate a query lets the clarification start earlier, when the clinician still remembers the encounter. It does not assign a code, it takes no clinical position, and it attacks the largest single source of delay.
The bottom rows are where demonstrations concentrate and where honest evaluation matters most. Complex surgical and inpatient work depends on judgment about clinical significance that frequently is not written in the document at all, which means no system reading the document can recover it. That is not a model capability gap that will close with a better model; it is missing information.
Testing a vendor's claim
Send these five questions, in writing, before any demonstration.
What was the accuracy measured against, and by how many coders? A single reviewer is not a reference standard. Ask for the adjudication procedure and the inter-reviewer disagreement rate on their own set.
What is the per-code-family breakdown, and what is the direction of the errors? An aggregate figure with no direction hides whether the system undercodes safely or overcodes expensively. Ask for both rates separately.
What happens on low confidence? The correct answer is that the chart routes to a coder with the ambiguity identified. A system that always produces an answer is a system that guesses on the hard cases, which are the cases you were paying it for.
What evidence does it show for each code? The specific text in the specific document that supports the assignment, viewable in one click. Without that, every review is a re-code from scratch and the system saves nothing. It is also what an auditor will ask for.
Will you run on our charts, blind, against our adjudicated set? A few hundred stratified charts, their output compared against your reference, scored by your people. Any vendor confident in their system will agree. This single test is worth more than every reference call, and it costs a fraction of a pilot.
Capacity without headcount
If the real problem is throughput rather than accuracy, four levers move it, roughly in order of return.
Reduce queries at the source. Group query causes, fix the top template and documentation defects. This removes work rather than redistributing it.
Triage by complexity. Route straightforward encounters to a fast path with automated suggestion and light review, and give the complex ones the time they need. Most organizations code everything through one process at one pace, which under-serves the hard charts and over-serves the easy ones.
Take non-coding work off coders. Chasing documents, chasing signatures, re-keying between systems, working denials that a billing edit should have caught. In most departments this is a meaningful fraction of the day and none of it requires a credential.
Automate narrowly and completely. One well-chosen domain running unattended, with a measured error rate and a confidence threshold that routes the rest to people, beats partial automation everywhere. Partial automation across the board leaves a human in the loop on every chart, which is where the time was in the first place.
Mistakes we see
- One accuracy percentage, covering four error modes with opposite consequences
- An audit standard set by one reviewer, so the measurement includes that reviewer's idiosyncrasies
- No direction reported, hiding whether the error is a lost dollar or an exposure
- Auditing coders when the defects originate in documentation templates
- Query data never grouped, so the largest source of delay is invisible
- Accepting a vendor figure measured on their evaluation set rather than your charts
- Automation with no confidence routing, which guesses on exactly the hard cases
- Code suggestions with no supporting text shown, so review costs as much as coding
When you do not need help with this
If your denial data shows structural errors — modifiers, bundling, laterality — the fix is edits in your billing system, configured by someone who knows your specialty. That is a configuration project, not a machine learning project, and it will return more per hour than anything else discussed here.
If you have never grouped your physician queries by cause, do that before talking to any vendor. It is a spreadsheet exercise on data you already have, it usually names three or four specific fixable defects, and it frequently removes enough rework to resolve the capacity pressure that started the search.
If your volume is modest and your specialty mix narrow, an experienced coder with good templates and a clean edit set is likely to outperform any automation you could buy, and to cost less. Automation earns its place at volume, in a bounded domain, with measurement around it.
Where outside engineering genuinely helps is building the evaluation infrastructure — the stratified set, the adjudication workflow, the per-family scoring with direction — and building the query-cause analysis on documents that are not queryable today. Those are real engineering problems and they are prerequisites for every decision above.
What a serious program has
- Four error modes measured separately, with direction reported
- Inter-coder agreement measured on your own charts, and treated as the ceiling
- A stratified, double-coded, adjudicated evaluation set that is refreshed
- Per-code-family scoring, never a single aggregate
- Query causes grouped by template, specialty and author, with owners
- Structural edits enabled and reviewed, with fire rates monitored
- Confidence routing that sends ambiguity to a person rather than guessing
- Supporting text shown for every suggested code, in one click
- Complexity-based triage instead of one path for every chart
- A stated position on the over-versus-under trade-off, decided by leadership rather than drifting
Bottom line
Coding accuracy is four measurements, not one, and two of them move in opposite directions when you tighten the rules. Before evaluating any system, measure how much your own qualified coders agree with each other, because that number is the ceiling on every claim you will be shown. Most of the defects start in documentation, so group the queries and fix the templates first — it is the one intervention that improves accuracy and capacity together with no trade-off. Then automate narrowly and completely in a bounded domain, with confidence routing, with the supporting text visible for every suggestion, and with per-family scoring that reports direction. And ask any vendor to run blind on your charts against your adjudicated set. The ones who agree are the ones worth talking to.
Frequently asked questions
It is real in narrow, templated domains where the source document is highly structured and the code space is small, and it has been running in some of those settings for years. It is not real for complex surgical or inpatient severity work, where the code depends on clinical judgment that is often not written in the document at all. Judge any claim by the domain it was measured in, not by the phrase.
They improve throughput and reduce the pressure that causes rushed work, which helps at the margin. They do not address documentation defects, template gaps, or the absence of structural edits, which is where most errors originate. If accuracy is the goal rather than backlog, the query analysis and the edit configuration will do more per hour than an additional coder.
Large enough that the per-code-family cells have real counts, which is the binding constraint and is what drives the size well past what people expect. A few hundred charts can support an overall figure and cannot support family-level conclusions. Stratify by complexity, size the strata to the questions you need answered, and be explicit that families with thin coverage were not measured.
That is a leadership decision, not a technical default, and it should be made explicitly and written down. Systematic undercoding has real costs, including distorted risk capture and a clinical picture that understates the population. The important thing is that the position is chosen and monitored rather than emerging accidentally from a threshold somebody set once.
Treat that as the answer. A blind run on a few hundred of your stratified charts against your adjudicated reference is a small effort for a vendor with a working system, and it is the only evidence that transfers to your specialty mix and documentation habits. Refusing it does not prove the system is weak, but it does mean the figure you were given is untested where it matters.
