The estimate you are asking for cannot be given honestly
A buyer asks four firms what it will cost to automate a review process. Two answer with a number and a timeline. One answers with a range so wide it is useless. One says the honest thing, which is that the answer depends on facts nobody in the conversation has yet — how consistent the source documents are, how much the experts agree with each other, what the current error rate actually is — and offers to go find those facts for a fraction of the build cost. The fourth firm usually loses the work, and is usually right. Everything below is about how to buy from the fourth firm without being the buyer who waits six months for certainty that is never going to arrive.

The core move is simple to state and hard to hold to: separate the question from the system, and buy the question first. You are not deciding whether to build software. You are deciding whether to spend a small, bounded amount finding out what is achievable, and then deciding about the software with real numbers in hand. Framed that way, technical uncertainty stops being a reason to stall and becomes a line item.
The bounded amount is smaller than most buyers expect. As a rule of thumb we would put ten to twenty percent of the projected build cost against removing the largest unknown, and we would want that work to take weeks rather than months. If a firm proposes a discovery phase that costs half the build and runs a quarter, that is not de-risking. That is the build with a different name on it.
You are probably here because
- You have been asked to approve a budget for something nobody can size
- One vendor promises 95 percent accuracy and another will not name a number at all
- A previous pilot produced a demo everybody liked and no decision anybody could act on
- You suspect the honest answer is “it depends” and you need a way to buy through that
The four uncertainties below are what “it depends” actually decomposes into, and each one has a cheap, decisive test.
Four uncertainties, and they are not equally dangerous
“We do not know if this will work” is four different statements wearing one coat. Pull them apart, because they have different tests, different costs, and very different consequences when ignored.
Data uncertainty. Does the information the system would need actually exist, in a form somebody can get to, in enough volume, covering the cases that matter? This is the one that ends projects and it is the cheapest to test. It is also the one buyers are most confident about and most often wrong about, because the person who says the data exists is rarely the person who has tried to pull it.
Ceiling uncertainty. Even done perfectly, how well can this task be performed? Not by a model — by anyone. Many tasks that sound crisp are genuinely ambiguous, and the ambiguity puts a hard lid on any achievable number.
Threshold uncertainty. How good does it have to be before it is worth deploying? This one is not technical at all, it is yours to answer, and it is astonishing how far into a project it can go unanswered.
Workflow uncertainty. If it worked, would the people whose work it changes actually use it, and would the change survive contact with the way they really do the job? This is where technically successful projects go to die, and no amount of model work touches it.
Which unknown actually kills the project — our ranking
How often each ends a project, as we rank it — judgment, not a study. The bottom row is the one that gets all the attention in the room.
That last row is the point of the chart. Buyers spend their anxiety on the modelling, which is the part with the most competent people working on it and the most mature tooling. The failures cluster at the top, in access, adoption and ambiguity, and those are questions you can answer without hiring anybody at all.
The cheapest decisive test for each one
For every uncertainty, ask what the smallest piece of work is that would change your decision. Not the work that would make you feel better — the work whose result would move you from build to do-not-build, or from one design to another. If the answer would not change what you do, do not buy it.
| Uncertainty | The test that resolves it | Rough cost | What a bad result means |
|---|---|---|---|
| Data exists and is reachable | Actually pull 500 real records end to end, through the real approval path | Days, mostly internal | Stop. Everything downstream was hypothetical |
| The task has a workable ceiling | Two experienced people label the same 100 items independently; measure how often they agree | A day of two people’s time | Redefine the task until people agree, before any modelling |
| Current performance | Measure today’s error rate on those same 100 items | A day | You may already be better than you thought, or far worse |
| The method can reach the ceiling | A time-boxed technical spike on the real sample, scored against the labels | 2–4 weeks of a small team | Either a different method, or a smaller slice of the problem |
| People will use it | Put the output in front of the people who would act on it, on real cases | Days, inside the spike | Change the surface, not the model. Or stop |
| It fits the systems you own | Write and run one real call against the system of record | Days | Budget for the integration, which is now the project |
Three of the six cost you almost nothing and need no vendor. The pattern in practice is that a buyer spends six weeks selecting a firm to answer questions the buyer could have answered in four days with two of their own people and a spreadsheet. Do those first. They also make you a far better customer, because you arrive at the vendor conversation with numbers instead of hopes.
Find the ceiling before you set the target
This is the most useful hour in the whole process and it is almost never spent. Take a hundred real cases. Have two of your experienced people work them independently, without conferring. Then count how often they agree.
Whatever that number is, it is approximately your ceiling. If two of your best people disagree on eighteen of a hundred cases, a target of ninety-five percent is not available to anyone, at any price, with any technology — because there is no stable answer to be right about on the cases where your own experts split. A system trained on labels those people produced will inherit their disagreement, and a system evaluated against those labels will be marked wrong on cases where it was arguably right.
A low agreement number is not a failure. It is the most valuable finding of the entire scoping exercise, and it usually points somewhere productive. Look at where the two people split. Often the task is really two tasks: a large easy majority where they agree completely, and a small hard tail where the answer is genuinely contested. Automate the first, route the second to a person, and you have a project with a defensible target instead of an argument waiting to happen at acceptance.
Do the same exercise on the current process while you have the cases open. Almost nobody knows their existing error rate. Without it, any number a vendor quotes is unreadable: ninety-one percent is a triumph against a manual process running at eighty-two and a disaster against one running at ninety-six.
Set the threshold from the cost of being wrong
Targets get set at round numbers because round numbers feel like standards. Ninety-five percent is not a standard. It is a number somebody said in a meeting, and it has ended more projects at acceptance than any technical problem.
Derive it instead. What does one wrong answer cost, in money, in time, in risk? What does a missed case cost? Those two are almost never equal, and once you write them down the target stops being a number and becomes an operating point. A system that flags twice as many cases as it needs to but misses almost nothing is exactly right when a miss is expensive and a false alarm costs an analyst thirty seconds. The same system is unusable when the flag triggers a customer letter.
Write the acceptance sentence at scoping time, before anyone builds anything: the metric, the population it is measured over, the threshold, and what happens if it is missed. If you cannot write that sentence, you are not ready to buy the build — but you are ready to buy the work that produces the sentence, which is a perfectly good thing to purchase.
Hold the evaluation set yourself
Whatever else you agree, keep the test data. You select it, you hold it, the vendor does not see it before the scoring run. This is not about trust. Ordinary development pressure is enough to drift a self-reported number upward: a hard batch gets called unrepresentative, an ambiguous label gets corrected, and the reported figure climbs while real performance stays flat. Two sentences in the scope prevent the entire class of argument, and no good firm will object to them.
Not sure which unknown is the expensive one?
Describe the problem in a paragraph or two to contact@precisionfederal.com — what happens today, what you would like to be different, and what data exists. You get back a written note naming the unknown we would resolve first and the cheapest way we know to resolve it, including the parts you can do yourselves without hiring anyone. One business day.
contact@precisionfederal.comTime-box the work, do not scope-box it
When the outcome is uncertain, fixing the scope and letting the time float is the wrong trade, because nobody can size work whose difficulty is unknown. Fix the time and the money instead, and let the finding vary.
The contract shape that works: a fixed fee, a fixed number of weeks, a named team, a defined dataset, and a specified set of questions to be answered in writing at the end. The deliverable is not a system. It is a document with numbers in it, the code and scripts that produced those numbers, and a recommendation the firm is willing to put its name on — including a recommendation not to proceed.
Four to six weeks is the range we would default to. Under three weeks there is not enough time to get through the data access, which is where the first two weeks usually go. Past eight, the work stops being a question and becomes an unmanaged build, and the findings arrive after the decision needed them.
Write the kill criteria into that agreement before it starts, while everyone is calm and nobody is invested. If the extract cannot be produced by week two. If expert agreement comes in under a level you name. If the spike cannot clear a floor you name on the held-out sample. Stopping at a written criterion is a decision. Stopping without one is a failure, politically, no matter how much money it saved — and that asymmetry is exactly why so many projects that should have stopped kept going.
Cost of learning it late instead of early — relative
Relative pain of a late discovery versus the same discovery in week two. Our ranking, not a measurement.
Scope the smallest slice that is still real
The instinct under uncertainty is to narrow, and narrowing is right. The failure is narrowing along the wrong axis. Cutting quality — a demo on clean data, a prototype that ignores the hard cases — produces something that proves nothing, because the hard cases are the entire question.
Cut breadth and keep depth. One document type instead of nine, but the messy real examples of that one type, all the way through to the place the answer has to land. One region, one product line, one customer segment. The test of a good slice is whether a decision follows from the result: if it works on this slice, does the same approach obviously extend, and does someone actually benefit today?
Include the awkward parts on purpose. The scanned pages. The records with missing fields. The month where somebody changed the process and nobody wrote it down. A slice that excludes them is a slice that has excluded the reason you needed help.
What a good scoping deliverable contains
At the end of four to six weeks, you should be holding a document you could hand to a different firm, a board, or a finance committee. Not a slide deck of possibilities.
The measured baseline. What the current process actually does, on real cases, with the sample and method stated.
The ceiling estimate. Expert agreement, where they disagreed, and what that implies about the maximum defensible target.
What the spike achieved, on data the vendor did not choose, reported against the same cases and broken down by segment — because a single average hides the segment where it fails completely.
What it cost to run at the volumes you named, in compute and in human review, per item and per month. Cost per item is the number that decides whether a working system is worth operating, and it is left out of most pilot reports.
The failure analysis. Which cases it got wrong, grouped by cause, in enough detail that you can judge whether those failures are tolerable in your process.
A recommendation with a number attached, and the conditions under which it changes. Including, where warranted, do not build this.
The mistakes we see most often
- Approving a build budget before anyone has pulled a real record
- A ninety-five percent target nobody derived from the cost of being wrong
- Measuring against labels made by one person, so the ceiling was never known
- A demo on clean data that answered a question nobody had
- The vendor holding the test set, making the acceptance number a self-report
- A discovery phase costing half the build and running a quarter — that is the build
- No kill criteria, so stopping became a political act rather than a planned one
- No cost per item, so a system that works turns out not to be worth running
Six weeks from “we should do something” to a fundable decision
Scoping arc
Weeks one and two are yours and cost you almost nothing but attention. They are also the two weeks that most often end the project, which is the best money you will ever not spend.
When the answer is no
Sometimes the honest finding is that the data does not support it, the ceiling is too low, or the economics do not work at your volumes. A firm that tells you this in week five has saved you the build, and has given you something specific: what would have to change for the answer to become yes. Often that is a data-capture change worth making on its own merits, and it turns a dead project into a two-year plan with a first step.
Treat a well-argued no as a delivered result, and pay for it without friction. Buyers who do this get honest findings for the rest of their careers. Buyers who punish it get optimism, and optimism is expensive at build scale.
Bottom line
You are not being asked to estimate a project. You are being asked to estimate a project whose difficulty is unknown, which is a different task with a different method. Split the unknown into its four parts, run the cheap decisive test for each — three of which need no vendor at all — find the ceiling from your own experts before you set any target, fix the time and money rather than the scope, and write the stopping rule while everybody is still calm. Do that and the estimate arrives with real numbers behind it in six weeks. Skip it and you will get an estimate in six days, and it will be a guess wearing a spreadsheet.
Frequently asked questions
As a rule of thumb, ten to twenty percent of the projected build, over four to six weeks. The test is not the percentage though — it is whether the deliverable is a decision. If the output is a written finding with measured numbers, a cost per item and a recommendation the firm will stand behind, it is scoping. If the output is a partly built system, it is the build with a different name and it should be priced and governed as one.
Derive it rather than choose it. Establish the ceiling by measuring how often two of your own experts agree on the same cases, measure what your current process achieves, then set a threshold from the cost of a wrong answer versus the cost of a missed one. Those two costs are rarely equal, and once they are written down the target usually stops being a single number and becomes an operating point with two figures.
They can, and the price will carry their uncertainty as a risk premium, or it will be low and become a change-order conversation. Neither is dishonest; both are expensive. The better shape is a fixed price for the scoping work, where the unknowns are bounded by time rather than by scope, followed by a fixed price for the build once the numbers exist. Firms are usually happy to commit hard to the second once they have done the first.
Small in breadth, never in difficulty. One document type, one region, one product line is fine. Clean data, curated examples and the easy cases removed is not, because the hard cases are the reason the project exists. The test of a slice is whether a real decision follows from the result and whether somebody actually benefits if it works.
Three things, and they take days rather than weeks. Pull real records through the real approval path, so data access is proven rather than assumed. Have two experienced people label the same hundred cases independently and measure their agreement. Measure the current process error rate on those same cases. Arriving at a vendor conversation with those three numbers changes the quality of every proposal you receive.
