The problem is not that you lack expertise
Buyers assume they are at a disadvantage because they do not understand the technology. That is rarely the binding constraint. Plenty of technically strong teams buy badly, and plenty of non-technical operators buy well. The difference is not knowledge of transformers. It is whether the buyer insisted on evidence that could have come out the other way. A claim you cannot imagine failing is not evidence. It is a sentence. Almost everything below is a way of turning sentences into things that could fail.

Two other things are worth saying up front, because they set expectations for the rest. First, most AI vendors are not lying. The claims in the deck are usually true in some measurable sense and simply do not describe your situation, which is a different and much more common failure. Second, the questions that work are boring. They are about denominators, test sets, escalation paths and export formats. Nobody in the room will find them impressive. They are the ones that change outcomes.
Four kinds of claim, and only two are worth your time
Before responding to anything, sort it. Every sentence in a vendor deck is one of four kinds, and they behave completely differently under pressure.
Capability claims. "Our system reads invoices." "We support your document types." Almost always true, almost never useful. Any modern system can produce output for any input. The claim contains no statement about quality, so there is nothing to check. Do not spend meeting time here.
Performance claims. "94% accuracy." "Cuts review time by half." These are checkable, and the number by itself means nothing at all until you know three other things. This is where most of your attention belongs.
Comparative claims. "Three times more accurate than the leading alternative." Checkable only if the vendor names the alternative, the workload and the version. If they will not name it, treat the sentence as decoration and move on without argument.
Outcome claims. "Saved a customer two million dollars a year." The least checkable and the most persuasive, which is a bad combination. A savings number is the output of a model of a business you cannot see, built by someone with an interest in the answer.
You are probably here because
- Three vendors quoted three different things and you cannot line them up
- The demo was excellent and you have no idea whether that means anything
- Someone in the room asked what happens when it gets it wrong and the answer was vague
- You have been asked to sign, and the only real evidence you have is that everyone was likeable
All four have the same fix, and it takes about three weeks: build a small test set of your own real work, run it past everyone identically, and read the errors rather than the averages.
A percentage without a denominator is a decoration
When a vendor says 94%, three questions turn that into information, and they can be asked in sequence in about ninety seconds.
What is the unit? Per character, per field, per page, per document, per case. This is not pedantry, it is arithmetic. A document-extraction system at 94% per field, over a form with thirty fields, has roughly a 16% chance of producing a document with every field correct, because 0.94 to the thirtieth power is about 0.16. Both numbers are honest. Only one of them describes a workflow where somebody needs a clean record at the end. Ask which unit your process actually consumes, then ask for the number in that unit.
What was it measured on? How many items, drawn from where, collected when, and how similar to your material. A number from a curated public corpus of clean scans tells you almost nothing about a queue of faxed and photographed documents. Ask whether the test set includes the cases your staff currently argue about, because those are the ones you are paying to fix.
Who decided what was correct? Ground truth is produced by people, and people disagree. Ask how disagreements were resolved and how often they occurred. If two of your own experts agree only 90% of the time on a judgment, no system can be scored above 90% against either of them, and a vendor claiming 97% on that task has measured something other than what you think.
There is one more trap worth naming because it recurs. In any task where the interesting cases are rare, plain accuracy is meaningless. If 0.6% of transactions are fraudulent, a system that answers "not fraud" every single time scores 99.4%. Ask instead for precision and recall at the operating point you would actually run, and ask what happens to both if you tighten the threshold. A vendor who cannot produce that curve has probably never had to run the system under a real workload.
The demo is a selected sample, and everyone in the room knows it
A demo is a performance, and there is nothing wrong with that. The way to get information out of one is to change the inputs. Bring five examples of your own work: three ordinary ones and two that your own people disagree about. Ask to run them live.
The reaction is the data. A vendor with a working, general system usually says yes, and it takes ten minutes. A vendor whose system needs per-customer configuration will say they can set that up for next week. Neither answer disqualifies anyone. They describe two very different products with two very different prices and timelines, and it is far better to learn which one you are buying in the first meeting than in month four.
One discipline: do not hand over your five cleanest examples. The instinct is to send tidy material so the test is fair. A fair test is not what you want. You want the test that resembles Tuesday.
| When they say | Ask | A strong answer sounds like | A weak answer sounds like |
|---|---|---|---|
| "94% accurate" | Per what, on how many, judged by whom? | "Per field, on 1,200 held-out documents, two annotators, 4% disagreement" | "That's from our internal benchmarking" |
| "State of the art" | On which benchmark, which version, when? | A named public leaderboard and a date | "Across the board" |
| "Works with your data" | Can we run five of ours now? | "Send them over, we will screen-share" | "We would need an onboarding phase first" |
| "Used by [large company]" | Which team, doing what, for how long? | A named function, a start date, an offer to connect you | "We're not able to share details" |
| "Saved them $2M" | Measured against what baseline, by whom? | A before number, a method, and what the customer stopped doing | "That's their figure" |
| "Enterprise-grade security" | Show me the data-handling clause in the contract | A clause reference and a sub-processor list | A link to the trust page |
What a piece of evidence is actually worth
Not all evidence is the same size, and the cheap kinds are the ones offered first. Our rough ordering, which is judgment rather than measurement, but the ordering is the useful part.
How much a form of evidence tells you
Our weighting, not a measurement. The gap between the top two and the bottom two is the point.
Logos are not references, and references are not testimonials
A logo means someone at that organization signed something at some point. It might have been a departmental trial that ended eighteen months ago. Ask which team, what they run today, since when, and whether you can speak to them. A vendor who names a specific function and a start date is telling you something. A vendor who cannot is telling you something too.
The single highest-yield hour in any evaluation is a reference call with no vendor on the line. Ask for it explicitly and expect mild resistance; a vendor confident in the relationship will arrange it. Three questions do most of the work. What surprised you after you signed? What did you end up doing yourselves that you expected them to do? If you were buying again tomorrow, what would you write differently into the contract? People answer that last one honestly in a way they never answer "would you recommend them."
Public benchmarks, and the small one you build yourself
Public benchmarks are useful for one thing: telling you whether a system is in the right league at all. They are weak evidence about your workload for two reasons. They are built from generic material, and they leak. Datasets published on the open internet end up in training data, and a strong score on a widely available benchmark cannot be cleanly separated from having seen it. Nobody can fully quantify that, including the vendor.
The test you build yourself has none of those problems. It is also cheaper than people expect. In our experience the shape that works is 150 to 300 real examples, drawn from actual work rather than a clean folder, labelled by the person who does the job today, split into an ordinary set and a hard set. That is typically two to four days of one employee's time. It is usually the least expensive line in the whole purchase, it does not expire, and it is the only artifact in the process that keeps its value if you change vendors or revisit the decision in two years.
Run the identical set past every vendor, with identical instructions, and score it yourself. The scoring matters more than the running. If you send the set out and let each vendor report their own result, you have learned about their optimism, not their systems.
Ask what happens when it is wrong
Every one of these systems is wrong sometimes, and the interesting question is never the average. It is the error behaviour. Three questions:
Does it know when it is unsure? A confidence signal is only useful if it is calibrated, meaning that things it scores at 0.7 are right about 70% of the time. Ask whether that has been checked, and against what. An uncalibrated score dressed up as a probability is worse than no score, because people build thresholds on it.
What does it do with a case it cannot handle? Refuse, escalate to a person, or answer anyway. A system at 88% accuracy that reliably routes its uncertain 12% to a human is usually worth more than a 94% system that is confidently wrong 6% of the time and silent about it. The second one moves your error rate from the queue into your customer relationships, where it is far more expensive.
What kinds of things does it get wrong? Not how often. What kind. A vendor who has looked will tell you: handwriting on carbon copies, documents over forty pages, one particular form revision, anything in a second language. A vendor who answers only with a percentage may not have examined the failures at all, and that is the more informative response.
Fifty errors, not one accuracy figure
Request fifty actual errors from a real evaluation: the input, what the system produced, what was correct. Ten minutes with that file tells you more than a quarter of benchmark numbers. You will see immediately whether the mistakes are small and recoverable or the kind that a person downstream would never catch. If a vendor cannot produce fifty errors, they have not run the system at enough volume to know how it behaves.
Outcome claims and the missing comparison
When a vendor says a customer saved two million dollars, the number is the output of a calculation you cannot see. Four questions make it either useful or clearly not: over what period, against what baseline, computed by whom, and what else changed in that period.
Most reported savings are a before-and-after comparison inside an organization that also reorganized, hired, retired a legacy system or changed a process in the same window. That does not make the figure a lie; it makes it uninformative about you. The question that cuts through it is simple and rarely asked: what did the customer stop doing? If a team was reassigned, a contract was cancelled or a queue was closed, there is a real saving with a name on it. If nothing stopped, the saving is projected capacity, which is a genuine benefit and a completely different kind of number.
Three facts about the system that you do need
You do not need to evaluate the architecture. You do need three facts, and all three belong in the contract rather than in a conversation.
Whose model is underneath, and can it change without telling you? A large share of AI products are an application layer over a third-party model. That is a reasonable engineering choice and often the right one. It also means the behaviour of what you bought can shift when the provider updates, without a line of the vendor's code changing. Ask what their policy is, whether they pin versions, and who bears the cost of a regression.
What happens to your data? Retention period, whether it is used to train anything, which sub-processors touch it, which region it sits in, and what happens on termination. Get this from the contract, not the security page on the website, and read whether the clause survives a change of ownership.
Can you get out? Export of your data, your labels, your configuration and anything you have tuned, in a documented format. Ask for a sample export file during evaluation. Requesting one at renewal, when leverage has moved, is how organizations discover their five years of corrections are not portable.
The mistakes we watch buyers make
- Judging on a demo built from the vendor's own examples, which is a test of their sales engineering
- Accepting an accuracy figure with no unit, and therefore never learning it was per field on a thirty-field form
- Skipping the unaccompanied reference call because it felt awkward to ask for
- Treating a successful pilot as evidence when the pilot was run by the vendor's best engineer, full time, on their best data
- Never asking what the system does with a case it cannot handle, and discovering the answer is "answers anyway"
- Comparing three quotes that priced three different scopes, and picking the cheapest description of the least work
- Signing without seeing an export file
A three-week evaluation that actually decides something
Vendor Evaluation
Step one is the one people skip and it is the one that pays. Writing down, before you buy, the specific number you expect to see at ninety days is the only mechanism that lets anyone tell afterwards whether the purchase worked. Without it, every outcome is a success, because success was never defined. Three sentences on a page, dated, signed by whoever owns the budget.
Before you sign
- Every performance number has a unit, a sample size and a source
- Your own examples have been run, live, on material you did not clean
- You have read fifty real errors and know what kind they are
- You know what the system does with a case it cannot handle
- At least two references spoke to you without the vendor on the call
- Data retention, training use and sub-processors are in the contract
- You have seen a sample export of your own data and configuration
- Someone owns a written expectation for day ninety
Bottom line
Almost every bad AI purchase we are asked to look at afterwards has the same shape. The evidence was all of a kind that could not have come out badly: a rehearsed demo, an unqualified percentage, a logo wall, a savings figure with no baseline. Nobody lied. Nobody checked. Replacing any two of those with something falsifiable, most cheaply your own small test set, changes the odds more than any amount of technical study. The vendor who is good will not mind. That, by itself, is the most reliable signal in the process.
Frequently asked questions
By itself, nothing. It becomes information once you know the unit of measurement, the size and origin of the test set, and who decided what counted as correct. The most common surprise is unit mismatch: a strong per-field number can translate into a weak per-document number, and your workflow usually consumes documents.
For a buying decision between vendors, 150 to 300 real examples is usually enough to separate them, and it takes two to four days of one knowledgeable employee's time. Split it into ordinary cases and hard cases and report the two separately, because the hard set is where vendors diverge and the ordinary set is where they all look similar.
Treat them as evidence a system is in the right league, not as a prediction about your work. Public datasets circulate widely enough that strong scores on them cannot be cleanly separated from familiarity with the data. A held-out set built after the model, or one you build yourself, is worth considerably more.
Ask what surprised them after signing, what work they ended up doing themselves that they had expected the vendor to do, and what they would write differently into the contract if they were buying again tomorrow. Ask for the call without the vendor present; the difference in candour is large.
Only if it is run the way production would be run. A pilot staffed by the vendor's strongest engineer, working full time on data they helped select, measures the ceiling rather than the expectation. Define in advance who does the work, what data is used, and what number would count as a pass.
