Skip to main content
Pilots & Proof of Value

Running a paid pilot that answers a real question

Most pilots end with a demonstration, a round of congratulations, and no decision. The cause is almost always the same: nobody wrote down the question at the start, so any result can be read as encouraging. Here is how to structure one so that it produces an answer you can act on, including the answer you did not want.

What this is A practical structure for a paid pilot, written by a firm that runs them. We benefit when you buy one, so the parts worth most to you are the section on what a pilot cannot prove and the section on doing it yourself with a spreadsheet. Both are honest.

The pilot that decides nothing

Here is the shape it takes. A few weeks of work, a meeting with a screen share, a model that handles the examples in the room, and a group of people who are pleased. Then somebody asks what happens next, and the room discovers that nobody knows what result would have meant yes and what result would have meant no. The conversation moves to a follow-up pilot, or to a larger version of the same thing, and six months later the process it was meant to change is still being done the old way.

Nothing went wrong technically. The failure was in the setup. A pilot is an experiment, and an experiment without a stated hypothesis and a stated threshold cannot fail, which also means it cannot succeed. All it can do is produce impressions.

Write the question as one sentence

Before any money is committed, someone has to write a sentence with three parts: a measurement, a threshold, and a decision that follows.

“If the system correctly extracts and matches at least 85% of last quarter's invoices, with under two seconds of processing each and under $0.03 per document, we will route the Midwest region through it in Q1 and reduce manual review to spot checks.”

That sentence does several jobs at once. It says what is being measured and on what data. It states the bar. It names what happens if the bar is met, which forces the person who would have to act to be in the room at the beginning rather than at the end. And it implies the opposite: if the number comes back at 61%, the region does not get routed, and everyone agreed to that in advance.

Weak versions of this sentence are easy to spot. “We want to see whether AI can help with invoices” has no threshold and no decision. “We want to explore what is possible” is a research budget, which is a legitimate purchase but a different one; call it that so the expectations match. The tell is whether the sentence could come back false. If it could not, you are buying a demonstration.

A pilot that can only succeed is not a pilot. If there is no result that would make you stop, you have already decided, and you are paying for a reason to feel confident about it.

Measure the baseline first, and do it before anyone builds anything

This is the step that gets skipped most often and costs the most when it is. You cannot show improvement against a number you never took.

Baseline means the current process, measured on the same units you will measure the new one on. How long does one item take, end to end, including the waiting. How many are done a week. How often is the result wrong, and how is wrong discovered. What does an hour of the person doing it cost, fully loaded. What happens downstream when there is an error — a credit note, an apology, a lost customer, a re-run.

Two or three days of watching and counting produces this, and it is worth doing whether or not you ever buy a pilot. It answers the prior question — is this worth automating at all — more reliably than any vendor conversation. Plenty of processes that feel painful turn out to consume six hours a month, and plenty of processes nobody complains about turn out to consume two days a week.

The baseline also gives you the honest denominator later. A model at 85% accuracy sounds mediocre until you measure the humans doing it now and find they are at 91% with a full day of latency, at which point 85% instantly and consistently, with the 15% routed to a person, is a different proposition entirely. Without the baseline that comparison is unavailable and the conversation turns into opinion.

Real data, a frozen holdout, and enough of it

A pilot on invented examples proves that the technology works on invented examples. That was not in doubt.

Real data. Your own records, with their real inconsistencies: the scanned pages, the supplier who changed their template in March, the entries where somebody typed the date in the notes field. Getting this released is often the longest pole in the whole engagement, which is why it belongs in week one and not week three.

A frozen holdout. Set aside a portion of the data before work begins, and do not let it be looked at until the final measurement. This is the single most effective protection against a pilot that looks better than it is. Without it, the natural rhythm of the work — try, look, adjust, try again — quietly tunes the system to the examples it has seen, and the final number describes memory rather than capability.

Enough of it. Twenty examples cannot distinguish 80% from 90%; the noise is larger than the difference you care about. A few hundred is usually enough to be worth arguing about, and a thousand is comfortable for most business processes. The right way to feel this is to ask what the range around the result is, not just the point estimate.

The hard cases on purpose. Ask the people who do the work to hand over their worst twenty examples. They will enjoy it, and those twenty tell you more about production behaviour than a thousand ordinary ones.

What to measure besides accuracy

Accuracy is the number everyone brings to the meeting and it is rarely the number that decides whether the thing survives contact with the business.

MeasureWhy it decides things
Review load createdIf 30% of items need a human check, you have moved the work rather than removed it. Compare the review minute count against the baseline minute count
Behaviour when wrongA system that is confidently wrong is worse than one that abstains. Ask what fraction of errors the system flagged itself
Cost at real volumePer-item cost multiplied by the annual count, plus infrastructure. Pilot volumes hide this and the multiplication is often the surprise
Latency where it mattersBatch overnight and answer-in-two-seconds are different systems at different prices. Decide which you need before, not after
Who is accountable for an errorAn unanswered version of this question stops more deployments than any technical result
What the people doing the work sayTwo of them, using it for a week, will tell you things no metric will. Their objections are usually specific and usually fixable

You are probably here because

  • A previous pilot went well and nothing changed afterwards
  • You have been offered a free pilot and are trying to work out the catch
  • Your board wants a number by the end of the quarter and you need the pilot to produce one
  • Two vendors are proposing pilots and the proposals are not comparable

The one-sentence question and the baseline section are the two things that make a pilot decidable. The rest is structure around them.

The three outcomes, agreed in advance

Write all three into the agreement, with the threshold attached to each. This costs nothing and changes the psychology of the final meeting completely, because nobody has to be the person who brings up stopping.

Proceed. The number cleared the bar. Name what happens next and roughly what it costs, so that the pilot ends with a decision rather than a new evaluation.

Stop. The number missed by enough that no reasonable amount of further work closes it, or the cost per item at volume makes the value negative. This is a good outcome. A stop for $60,000 in eight weeks has protected a build that would have cost several times that and consumed a year of belief.

Redesign the question. The middle case, and more common than either. The result missed the bar but the failures cluster in a way that suggests a narrower scope works — one document type, one region, one customer segment. This outcome is only available if somebody looked at the errors rather than the average, which is why error analysis should be a named deliverable and not a courtesy.

Attach a kill criterion with a date to the middle of the pilot as well. “If we do not have production data access by day ten, we pause and reprice.” Access delays are the most common way a six-week pilot becomes a two-week pilot with four weeks of billing attached, and both sides are better off naming it early.

Duration, price, and who does what

Ranges vary with data access and problem shape, but these are the bands most work falls into.

ShapeDurationTypical rangeAnswers
Narrow feasibility test1–2 weeks$8,000–$25,000One question on existing data. No integration, no interface
Standard pilot4–8 weeks$30,000–$90,000Real data end to end, measured against baseline, a rough interface for the reviewers
Pilot with limited live use8–12 weeks$90,000–$200,000The above plus a small group using it on live work, which is where adoption evidence comes from

Your own cost is real and is usually underestimated. Expect to spend, on your side: someone senior for two to four hours a week, the person who knows the process for perhaps a day a week during the first fortnight, whatever it takes to get data access approved, and a few days of somebody's time labelling or checking examples. If nobody on your side has that time, the pilot will slip regardless of the vendor, and it is better to delay the start than to run it thin.

On free pilots: they exist, and the trade is visible if you look for it. A free pilot is a sales expense, so it is staffed accordingly, scoped to what demonstrates well, and comes with an expectation. Paid pilots get the people who would do the build, and they let you say no at the end without owing anybody anything. If a free pilot is offered, ask who specifically will do the work and whether the same people continue into the build.

What a well-run pilot actually settles — our read

Whether the data supports the task
94
Quality reachable at this scope
86
Cost per item at real volume
74
Effort to reach production
52
Whether staff will adopt it
34
Long-run maintenance burden
22

Our judgment, not a study. The bottom two rows are what pilots are most often assumed to settle and least able to.

What a pilot cannot prove

Feasibility is not adoption. A pilot can show that a system produces good answers on your data. It cannot show that the team will change how they work, that the exception process holds up on a bad Monday, or that the thing survives the quarter when the person who championed it goes on leave. That evidence comes only from live use by people who did not choose the project.

A pilot also underestimates the last mile, reliably and in the same direction. Retries, monitoring, permissions, audit trails, the admin screen someone needs when a record is wrong, the runbook for when the upstream system changes a field. As a planning rule, treat the pilot as roughly a quarter to a third of the work to a dependable production system, and budget the rest before you celebrate. The pilot is the cheap part precisely because it skips all of that on purpose.

And a pilot proves nothing about the second use case. Systems that handle one document type well often need substantial rework for the next one. Ask what would change if you added the second type, and treat a confident “nothing” with mild suspicion.

Why pilots stall

  • Data access takes five weeks of the six — and nobody named an owner for it on day one
  • The success criteria get negotiated after the results arrive, which is not measurement
  • No baseline, so the outcome cannot be compared to anything and becomes an opinion
  • The decision-maker is not in the room until the final presentation
  • The pilot runs on curated examples, and the curation is discovered afterwards
  • Only averages are reported, so nobody looked at where the failures cluster
  • No named end date, so it drifts into a small standing project nobody decides about
  • Nothing was written about what the team does differently on the Monday after

When you should not buy one

Three cases, said plainly.

The question is answerable in a spreadsheet. If two hundred rows and an afternoon would tell you whether the pattern exists, do that first. You will either save the whole cost of the pilot or walk into it knowing considerably more than the vendor.

The value is too small to matter. If the baseline arithmetic says the process consumes $30,000 a year of effort, a $60,000 pilot in front of a $250,000 build is not a close call. This is why the baseline comes first: it sometimes ends the conversation for free.

The blocker is not technical. If legal has not decided what data may be processed, or two departments disagree about who owns the process, a pilot will produce a good number and change nothing. Go and settle that first. It costs no money and it is the actual critical path.

Before you sign a pilot

  • One sentence states the measurement, the threshold and the decision
  • The person who would act on a yes has read and agreed that sentence
  • The baseline is measured before the build starts
  • Real data, with a holdout frozen before work begins
  • A named owner for data access, with a date and a pause clause
  • Error analysis is a deliverable, not just an average
  • Review load, cost at volume and failure behaviour are measured alongside accuracy
  • All three outcomes — proceed, stop, redesign — are written down in advance
  • You keep the code, the evaluation set and the labels whatever the outcome
  • A fixed end date, and a rough cost for the production version if it proceeds

Bottom line

The difference between a pilot that decides something and one that does not is set in the first week, not the last. Write the question with a number and a decision in it. Measure what you do today before you measure the alternative. Use your own messy data, freeze a holdout, and look at the failures rather than the average. Agree in advance what result would make you stop, and say out loud that stopping is a permitted answer. Done that way, a pilot is one of the cheapest decisions available to you. Done the other way, it is an expensive demonstration of something nobody doubted.

Frequently asked questions

How long should a pilot run?

Four to eight weeks for most business processes, and one to two for a narrow feasibility check. Beyond eight weeks without live use, additional time tends to produce refinement rather than new information. If the schedule is being driven by waiting for data access, that is not pilot duration — it is a dependency that should have a date and a pause clause attached to it.

Should we accept a free pilot?

You can, with your eyes open. A free pilot is a sales expense, which means it is staffed with whoever is available, scoped toward what demonstrates well, and carries an unstated expectation of what comes next. Paying a modest amount buys the people who would do the build and a clean exit at the end. If you take a free one, ask who specifically does the work and whether they continue into the project.

What accuracy is good enough?

There is no universal number and anyone offering one is guessing. It depends on what the current process achieves, what an error costs you, and whether the system can flag its own uncertainty. Eighty-five percent with reliable self-flagging and a fast human review path often beats ninety-five percent that fails silently, because you can build a process around the first and cannot around the second.

Who should own the pilot internally?

The person who owns the process being changed, not the technology function. They can get data released, they can put two of their staff on the review, and they can act on the result without asking anyone. A pilot sponsored by technology and aimed at someone else's process tends to produce a good result that nobody adopts, which is the most expensive kind of success.

What do we keep if we stop?

Everything worth keeping, and it should be in the agreement: the code, the evaluation set, the labels, the baseline measurements and the written findings. The evaluation set in particular is a durable asset — it is what lets you compare any future attempt, by any vendor or your own team, against the same bar. A pilot that ends with a stop and leaves you those artifacts was still a good purchase.

1 business day response

Have a pilot proposal in front of you?

Send it and we will tell you whether it can produce a decision as written, and what we would add or remove. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
PilotsBaselinesEvaluationDecision Criteria