The accuracy number nobody believes
Every company selling a model has a number on a slide. It is usually between 94% and 99%, it has no denominator, no date, no population, and no confidence interval, and the person it is shown to has seen forty of them this year. The number does not persuade. It signals that a marketing department has been near the model, which starts the buyer's technical team looking for the trick rather than looking for the fit. The evaluation then takes eleven weeks instead of three, because everything after the slide is spent reconstructing what the number meant.
The alternative is not a lower number. It is a complete one. "Character-level error rate of 2.3% across 41,800 pages sampled from the 2025 filing corpus, rising to 7.9% on carbon-copy scans and 11.4% on handwriting in the remarks field, measured against a double-keyed reference set" is a harder sentence to write and a much easier sentence to sell. It survives contact with an engineer. It gives a buyer something to plan around instead of something to test.
This is not a moral argument about candor. It is an argument about where deal time goes and who ends up carrying the risk when a number is left vague.
Why the vague number costs you the deal you thought it won
A buyer's technical reviewer has one job during evaluation: find the conditions under which your system stops working. They will find them. Every model has a slice where performance falls apart, and a reviewer with two weeks and a sample of real data will locate it. The only variable you control is whether they find it in your document or on their own.
Finding it in your document costs you nothing. It reads as a known boundary, described by people who measured it, with a stated mitigation. Finding it on their own costs you the deal or the price, because it reframes everything else you said. The reviewer now assumes the rest of the claims were built the same way, and the assumption is usually correct. One discovered gap between a published number and a measured one turns a technical evaluation into an audit.
The cost shows up in the calendar before it shows up in the price. A commercial procurement that starts from a published methodology tends to close in one to two months of technical evaluation. The same procurement starting from a bare accuracy claim runs three to six, because the buyer has to build the evaluation you did not give them. At a $400,000 to $2M engagement, four months of delayed start is a real number on both sides of the table.
What a buyer's reviewer actually checks first
Editorial weighting from evaluation work and practitioner reading — illustrative, not a measured statistic.
The claim you print is a claim you can be held to
There is a second reason the vague number is expensive, and it has nothing to do with sales cycles. Published accuracy claims are enforcement surfaces.
In March 2024 the SEC settled charges against two investment advisers, Delphia (USA) Inc. and Global Predictions Inc., over statements about their use of AI that the orders found false or misleading. The firms paid $400,000 in combined civil penalties, $225,000 and $175,000. The agency called the conduct AI washing. That September the FTC announced Operation AI Comply, five actions against companies it alleged used AI claims to supercharge deceptive conduct, including a settlement in which DoNotPay agreed to pay $193,000 over marketing of a service the complaint said had not been tested against the standard advertised.
The penalties are small next to an enterprise contract. The exposure is not the fine. It is that a published claim your engineering cannot reproduce sits in your own archive forever, discoverable in a customer dispute, an acquisition, or a diligence process three years from now, attached to a company someone else will own by then. A number with its methodology printed underneath is defensible on its face. A number without one is a statement whose meaning your counterparty gets to define later.
What the frameworks already ask you to write down
Most of the work is specified somewhere already, which means you are not inventing a format and your buyer is not evaluating one.
The NIST AI Risk Management Framework organizes the material under its Measure function, which asks for identified metrics, documented evaluation methods, and characterization of performance across relevant conditions, and its Map function, which asks for intended purpose and out-of-scope uses. ISO/IEC 42001, the AI management system standard published in December 2023, wants the same evidence held as a managed record rather than a one-time artifact. For anything touching a financial institution, SR 11-7, the Federal Reserve and OCC guidance on model risk management from 2011, has required documented model limitations and outcomes analysis for fifteen years. Your banking customers have a validation group whose entire job is checking for it.
On the security side, the OWASP Top 10 for LLM Applications and MITRE ATLAS both describe failure modes an evaluation should cover, and a document that names the ones you tested for reads very differently from one that does not. If federal work is anywhere in your future, NIST SP 800-53 assessment procedures and the FedRAMP process both run on evidence rather than assertion, and a firm already writing evaluation reports has a much shorter path into that world.
None of these requires certification to be useful. Naming the framework you organized the document under tells the reader what to expect and where to look, which is most of what a structure is for.
The document itself, section by section
What follows is the form we build for clients. It runs eight to fifteen pages, gets published or handed out under NDA depending on the sales motion, and replaces about forty slides of back-and-forth.
Scope and intended use. What the system decides, who operates it, what it is not for. The out-of-scope list is the section buyers read twice, because it tells them whether you understand your own product.
The evaluation set. Where the data came from, how many items, how it was sampled, what date range it covers, and how it was kept out of training. If the buyer cannot tell whether your test data leaked into your training data, every number after this is unreadable.
Reference labels. Who produced ground truth, what their qualifications were, and what the agreement rate was between annotators. A model reported at 96% against labels two humans agreed on 91% of the time is telling you something about the labels, not the model.
Headline performance with an interval. The rate, the sample size, and a confidence interval. A 2% error rate on 300 items and a 2% error rate on 40,000 items are different claims, and the interval is what makes the difference visible without an argument.
Slice performance. The table that carries the document. Error rate broken out by every dimension that matters operationally: input quality, document type, language, geography, customer segment, time period. Include the worst slice. Especially include the worst slice.
Failure modes with examples. Three or four real cases where the system was wrong, with the input, the output, and the reason. Redact what you must. Nothing else you can put in the document builds as much confidence as showing that you have looked at your own errors closely enough to explain them.
Operating points. Precision and recall at two or three thresholds, so a buyer can pick the one their workflow needs instead of accepting yours. This section alone often decides the fit.
Monitoring and drift. How the number is kept current, what triggers a re-measure, and what the buyer will see if performance moves after deployment.
The slice table is the whole document
An aggregate error rate is an average over a population you chose, and averages hide exactly the thing a buyer needs to know. A document classifier at 3% overall error can be at 1% on clean digital PDFs, which are 80% of the corpus, and 19% on faxed scans, which are 20% of the corpus and 100% of one customer's volume. The aggregate is honest and useless. The slice table is what lets a buyer decide whether the system fits their data instead of somebody else's.
Building it is not hard, and the reason most vendors skip it is not difficulty. It is that the slice table names the customers you should not sell to. That is a feature. A deal closed into the 19% slice becomes a failed deployment, a bad reference, and a renewal you do not get. The slice table is a qualification instrument pretending to be a technical appendix.
| Claim as usually written | The same claim, complete | What the second one does |
|---|---|---|
| "99% accurate" | Error rate 1.1% on 22,400 held-out records, 95% CI 0.97 to 1.24 | Survives a statistician. Fixes the sample-size argument before it starts |
| "Works on all document types" | Under 2% on nine of eleven types; 6.8% on multi-column tables, 9.1% on handwriting | Qualifies the buyer's corpus in one reading instead of one pilot |
| "Validated against industry data" | Reference set double-keyed by two annotators, inter-annotator agreement 0.94 kappa | Puts a ceiling on the claim and shows you know where it is |
| "Continuously improving" | Re-measured quarterly on a rolling window; last four quarters printed | Turns a promise into a record the buyer can check next year |
| "Enterprise-grade reliability" | p95 latency 340ms at 200 rps, measured on the deployed configuration | Replaces an adjective with something the buyer's architect can size |
What it costs to build
For a company with a model already in production and evaluation data lying around in notebooks, a first published evaluation report is roughly a four to eight week effort. The engineering is two to four weeks: assembling a defensible held-out set, rebuilding the evaluation so it runs from a single command, computing slices and intervals, and pulling real failure examples. The writing and internal review is another two to four, and the review is the slow part because legal and sales both have to get comfortable with printing a real number.
In consulting terms that lands between $60,000 and $150,000 depending on how much of the evaluation infrastructure exists already. The recurring cost afterward is small, usually a few days a quarter, because the expensive part was making the measurement reproducible rather than making it once.
Compare that against the alternative you are already paying for: bespoke pilots. Most companies without a published evaluation run a custom proof of concept for every serious prospect, at three to eight engineering weeks each, with a close rate under half. Two avoided pilots pay for the document, and the document keeps working after the pilots would have ended.
Where the evaluation effort goes on a first published report
Typical ranges for a model already in production. Illustrative, not a quote.
The objections, and which ones are real
"Competitors will use our number against us." They will quote it, and the quote will include your methodology, because a number without one is not usable as ammunition. Meanwhile their own claim has no denominator, and every buyer who has read your document now knows to ask for one. Publishing shifts the standard of evidence in the direction of whoever went first.
"Our number is not good enough." Sometimes true, and worth knowing before a customer discovers it. More often the number is fine and the fear is about the comparison to competitors' fictional numbers. Those competitors are being evaluated by the same reviewers, and the reviewers know.
"Legal will not allow it." Legal is usually more comfortable with a measured, bounded, dated claim than with the marketing sentence already on the website. A published error rate with a stated methodology is far easier to defend than "industry-leading accuracy" if anyone ever asks what that meant.
"Our performance changes too fast to publish." Then publish the date and the version, and re-publish quarterly. A number stamped with its measurement date is honest. A number with no date is the one that ages badly, because it implies a currency it does not have.
How this changes the sales conversation
The mechanical effect is that the technical evaluation moves earlier and gets shorter. A buyer who has your evaluation report before the first meeting arrives with questions about fit rather than questions about credibility. The demo stops being an argument about whether the system works and becomes a conversation about which operating point suits their workflow. That is a different meeting with a different close rate.
The second effect is on price. Vendors who cannot describe their failure modes get priced as commodities, because the only comparable a buyer has is the accuracy number and everyone's accuracy number is 97%. A vendor who shows the slice table is selling a characterized system, and characterized systems get compared on fit rather than on price per seat.
The third effect is on the contract. Acceptance criteria written from a real evaluation are written once. Acceptance criteria written from a marketing claim get renegotiated during deployment, when the model meets the customer's actual data and the gap between the slide and the corpus becomes everyone's problem at the same time.
Where this lands for a firm evaluating engineering partners
The same test applies to whoever builds this for you. A firm that proposes an evaluation report should be able to describe the last one they built, what the worst slice was, and what the client decided to do about it. If the answer is abstract, the offer is a template.
The work is specific: getting a held-out set that is genuinely held out, deciding which slices matter for a given business, computing intervals correctly for rare events, and writing the failure section in language a customer's procurement team can read without an interpreter. It sits between data engineering, statistics, and technical writing, and the reason it is rarely done well is that most teams are strong in one of the three.
Bottom line
Publishing a measured error rate is not candor for its own sake. It moves the discovery of your worst case from the buyer's test set into your own document, where it costs nothing. It converts a marketing claim that creates enforcement exposure into a bounded technical statement that reduces it. It qualifies bad-fit prospects before they become failed deployments. And it shortens the part of the sales cycle that everybody hates, which is the eight weeks a technical reviewer spends reconstructing what your number meant. The number you are least comfortable printing is usually the one doing the most work once it is printed.
Frequently asked questions
Generally less than the alternative. A dated, bounded claim with a stated methodology is defensible on its face. An unqualified accuracy claim in marketing material is the shape of statement the SEC and FTC have pursued as AI washing, and it is the one your counterparty gets to interpret later.
Common practice is a public summary carrying the headline rate, the sample size, the date, and the main slice breakouts, with the full report including failure examples and per-threshold behavior released under NDA during evaluation. The public version has to contain enough to be checkable, or it reads as another slide.
Print it with the mitigation next to it: a confidence threshold that routes those cases to human review, a preprocessing step, or a stated exclusion from scope. A named boundary with a control attached reads as engineering. The same boundary discovered by the buyer reads as concealment.
Quarterly is the common cadence, plus any release that changes model behavior, training data, or intended use. The cost of a re-measure is small once the evaluation runs from a single command, which is the main reason to build it that way the first time.
NIST AI RMF Measure and Map subcategories are the most widely recognized starting point and cost nothing to adopt. Add ISO/IEC 42001 record-keeping structure if you are pursuing that certification, and SR 11-7 sections if your buyers are banks or insurers with a model validation group.