Skip to main content
Scoping & Buying

Scoping an AI project when you do not know what is possible

Nobody can tell you what accuracy is achievable on your data until somebody looks at your data. That is not a reason to wait. It is the first thing to buy, it is usually the cheapest part of the project, and there is a right way to buy it.

The estimate you are asking for cannot be given honestly

A buyer asks four firms what it will cost to automate a review process. Two answer with a number and a timeline. One answers with a range so wide it is useless. One says the honest thing, which is that the answer depends on facts nobody in the conversation has yet — how consistent the source documents are, how much the experts agree with each other, what the current error rate actually is — and offers to go find those facts for a fraction of the build cost. The fourth firm usually loses the work, and is usually right. Everything below is about how to buy from the fourth firm without being the buyer who waits six months for certainty that is never going to arrive.

The core move is simple to state and hard to hold to: separate the question from the system, and buy the question first. You are not deciding whether to build software. You are deciding whether to spend a small, bounded amount finding out what is achievable, and then deciding about the software with real numbers in hand. Framed that way, technical uncertainty stops being a reason to stall and becomes a line item.

The bounded amount is smaller than most buyers expect. As a rule of thumb we would put ten to twenty percent of the projected build cost against removing the largest unknown, and we would want that work to take weeks rather than months. If a firm proposes a discovery phase that costs half the build and runs a quarter, that is not de-risking. That is the build with a different name on it.

You are probably here because

  • You have been asked to approve a budget for something nobody can size
  • One vendor promises 95 percent accuracy and another will not name a number at all
  • A previous pilot produced a demo everybody liked and no decision anybody could act on
  • You suspect the honest answer is “it depends” and you need a way to buy through that

The four uncertainties below are what “it depends” actually decomposes into, and each one has a cheap, decisive test.

Four uncertainties, and they are not equally dangerous

“We do not know if this will work” is four different statements wearing one coat. Pull them apart, because they have different tests, different costs, and very different consequences when ignored.

Data uncertainty. Does the information the system would need actually exist, in a form somebody can get to, in enough volume, covering the cases that matter? This is the one that ends projects and it is the cheapest to test. It is also the one buyers are most confident about and most often wrong about, because the person who says the data exists is rarely the person who has tried to pull it.

Ceiling uncertainty. Even done perfectly, how well can this task be performed? Not by a model — by anyone. Many tasks that sound crisp are genuinely ambiguous, and the ambiguity puts a hard lid on any achievable number.

Threshold uncertainty. How good does it have to be before it is worth deploying? This one is not technical at all, it is yours to answer, and it is astonishing how far into a project it can go unanswered.

Workflow uncertainty. If it worked, would the people whose work it changes actually use it, and would the change survive contact with the way they really do the job? This is where technically successful projects go to die, and no amount of model work touches it.

Which unknown actually kills the project — our ranking

Data does not exist, or cannot be released to anyone
94
Nobody will change how they work to use it
86
The task is genuinely ambiguous, so the ceiling is low
75
No agreed threshold, so nothing can ever be declared done
68
Integration into the system of record is harder than expected
50
The modelling itself turns out to be hard
26

How often each ends a project, as we rank it — judgment, not a study. The bottom row is the one that gets all the attention in the room.

That last row is the point of the chart. Buyers spend their anxiety on the modelling, which is the part with the most competent people working on it and the most mature tooling. The failures cluster at the top, in access, adoption and ambiguity, and those are questions you can answer without hiring anybody at all.

The cheapest decisive test for each one

For every uncertainty, ask what the smallest piece of work is that would change your decision. Not the work that would make you feel better — the work whose result would move you from build to do-not-build, or from one design to another. If the answer would not change what you do, do not buy it.

UncertaintyThe test that resolves itRough costWhat a bad result means
Data exists and is reachableActually pull 500 real records end to end, through the real approval pathDays, mostly internalStop. Everything downstream was hypothetical
The task has a workable ceilingTwo experienced people label the same 100 items independently; measure how often they agreeA day of two people’s timeRedefine the task until people agree, before any modelling
Current performanceMeasure today’s error rate on those same 100 itemsA dayYou may already be better than you thought, or far worse
The method can reach the ceilingA time-boxed technical spike on the real sample, scored against the labels2–4 weeks of a small teamEither a different method, or a smaller slice of the problem
People will use itPut the output in front of the people who would act on it, on real casesDays, inside the spikeChange the surface, not the model. Or stop
It fits the systems you ownWrite and run one real call against the system of recordDaysBudget for the integration, which is now the project

Three of the six cost you almost nothing and need no vendor. The pattern in practice is that a buyer spends six weeks selecting a firm to answer questions the buyer could have answered in four days with two of their own people and a spreadsheet. Do those first. They also make you a far better customer, because you arrive at the vendor conversation with numbers instead of hopes.

Find the ceiling before you set the target

This is the most useful hour in the whole process and it is almost never spent. Take a hundred real cases. Have two of your experienced people work them independently, without conferring. Then count how often they agree.

Whatever that number is, it is approximately your ceiling. If two of your best people disagree on eighteen of a hundred cases, a target of ninety-five percent is not available to anyone, at any price, with any technology — because there is no stable answer to be right about on the cases where your own experts split. A system trained on labels those people produced will inherit their disagreement, and a system evaluated against those labels will be marked wrong on cases where it was arguably right.

If two of your best people disagree on eighteen cases out of a hundred, ninety-five percent is not available to anyone, at any price. The disagreement is the ceiling.

A low agreement number is not a failure. It is the most valuable finding of the entire scoping exercise, and it usually points somewhere productive. Look at where the two people split. Often the task is really two tasks: a large easy majority where they agree completely, and a small hard tail where the answer is genuinely contested. Automate the first, route the second to a person, and you have a project with a defensible target instead of an argument waiting to happen at acceptance.

Do the same exercise on the current process while you have the cases open. Almost nobody knows their existing error rate. Without it, any number a vendor quotes is unreadable: ninety-one percent is a triumph against a manual process running at eighty-two and a disaster against one running at ninety-six.

Set the threshold from the cost of being wrong

Targets get set at round numbers because round numbers feel like standards. Ninety-five percent is not a standard. It is a number somebody said in a meeting, and it has ended more projects at acceptance than any technical problem.

Derive it instead. What does one wrong answer cost, in money, in time, in risk? What does a missed case cost? Those two are almost never equal, and once you write them down the target stops being a number and becomes an operating point. A system that flags twice as many cases as it needs to but misses almost nothing is exactly right when a miss is expensive and a false alarm costs an analyst thirty seconds. The same system is unusable when the flag triggers a customer letter.

Write the acceptance sentence at scoping time, before anyone builds anything: the metric, the population it is measured over, the threshold, and what happens if it is missed. If you cannot write that sentence, you are not ready to buy the build — but you are ready to buy the work that produces the sentence, which is a perfectly good thing to purchase.

Scoping Note

Hold the evaluation set yourself

Whatever else you agree, keep the test data. You select it, you hold it, the vendor does not see it before the scoring run. This is not about trust. Ordinary development pressure is enough to drift a self-reported number upward: a hard batch gets called unrepresentative, an ambiguous label gets corrected, and the reported figure climbs while real performance stays flat. Two sentences in the scope prevent the entire class of argument, and no good firm will object to them.

Not sure which unknown is the expensive one?

Describe the problem in a paragraph or two to contact@precisionfederal.com — what happens today, what you would like to be different, and what data exists. You get back a written note naming the unknown we would resolve first and the cheapest way we know to resolve it, including the parts you can do yourselves without hiring anyone. One business day.

contact@precisionfederal.com

Time-box the work, do not scope-box it

When the outcome is uncertain, fixing the scope and letting the time float is the wrong trade, because nobody can size work whose difficulty is unknown. Fix the time and the money instead, and let the finding vary.

The contract shape that works: a fixed fee, a fixed number of weeks, a named team, a defined dataset, and a specified set of questions to be answered in writing at the end. The deliverable is not a system. It is a document with numbers in it, the code and scripts that produced those numbers, and a recommendation the firm is willing to put its name on — including a recommendation not to proceed.

Four to six weeks is the range we would default to. Under three weeks there is not enough time to get through the data access, which is where the first two weeks usually go. Past eight, the work stops being a question and becomes an unmanaged build, and the findings arrive after the decision needed them.

Write the kill criteria into that agreement before it starts, while everyone is calm and nobody is invested. If the extract cannot be produced by week two. If expert agreement comes in under a level you name. If the spike cannot clear a floor you name on the held-out sample. Stopping at a written criterion is a decision. Stopping without one is a failure, politically, no matter how much money it saved — and that asymmetry is exactly why so many projects that should have stopped kept going.

Cost of learning it late instead of early — relative

Data cannot be released — found after the build starts
96
Ceiling is below the promised target — found at acceptance
88
The team will not adopt it — found at rollout
80
Threshold was never agreed — found in the acceptance meeting
64
Integration is bigger than the model work — found in month three
46
A different method would have been better — found late
24

Relative pain of a late discovery versus the same discovery in week two. Our ranking, not a measurement.

Scope the smallest slice that is still real

The instinct under uncertainty is to narrow, and narrowing is right. The failure is narrowing along the wrong axis. Cutting quality — a demo on clean data, a prototype that ignores the hard cases — produces something that proves nothing, because the hard cases are the entire question.

Cut breadth and keep depth. One document type instead of nine, but the messy real examples of that one type, all the way through to the place the answer has to land. One region, one product line, one customer segment. The test of a good slice is whether a decision follows from the result: if it works on this slice, does the same approach obviously extend, and does someone actually benefit today?

Include the awkward parts on purpose. The scanned pages. The records with missing fields. The month where somebody changed the process and nobody wrote it down. A slice that excludes them is a slice that has excluded the reason you needed help.

Cut breadth, never depth. A prototype that skips the hard cases has skipped the question you were paying to answer.

What a good scoping deliverable contains

At the end of four to six weeks, you should be holding a document you could hand to a different firm, a board, or a finance committee. Not a slide deck of possibilities.

The measured baseline. What the current process actually does, on real cases, with the sample and method stated.

The ceiling estimate. Expert agreement, where they disagreed, and what that implies about the maximum defensible target.

What the spike achieved, on data the vendor did not choose, reported against the same cases and broken down by segment — because a single average hides the segment where it fails completely.

What it cost to run at the volumes you named, in compute and in human review, per item and per month. Cost per item is the number that decides whether a working system is worth operating, and it is left out of most pilot reports.

The failure analysis. Which cases it got wrong, grouped by cause, in enough detail that you can judge whether those failures are tolerable in your process.

A recommendation with a number attached, and the conditions under which it changes. Including, where warranted, do not build this.

The mistakes we see most often

  • Approving a build budget before anyone has pulled a real record
  • A ninety-five percent target nobody derived from the cost of being wrong
  • Measuring against labels made by one person, so the ceiling was never known
  • A demo on clean data that answered a question nobody had
  • The vendor holding the test set, making the acceptance number a self-report
  • A discovery phase costing half the build and running a quarter — that is the build
  • No kill criteria, so stopping became a political act rather than a planned one
  • No cost per item, so a system that works turns out not to be worth running

Six weeks from “we should do something” to a fundable decision

Scoping arc

1
Internal only: pull 500 real records through the real approval path. Nothing else proceeds until this works
Week 1
2
Two experts label 100 cases independently. Measure agreement and measure today’s error rate
Week 1–2
3
Write the acceptance sentence and the kill criteria from the cost of being wrong
Week 2
4
Time-boxed technical spike on the real sample, scored on cases you selected and hold
Weeks 3–5
5
Put the output in front of the people who would act on it, on live cases, and watch
Week 5
6
Written findings: baseline, ceiling, achieved, cost per item, failure analysis, recommendation
Week 6

Weeks one and two are yours and cost you almost nothing but attention. They are also the two weeks that most often end the project, which is the best money you will ever not spend.

When the answer is no

Sometimes the honest finding is that the data does not support it, the ceiling is too low, or the economics do not work at your volumes. A firm that tells you this in week five has saved you the build, and has given you something specific: what would have to change for the answer to become yes. Often that is a data-capture change worth making on its own merits, and it turns a dead project into a two-year plan with a first step.

Treat a well-argued no as a delivered result, and pay for it without friction. Buyers who do this get honest findings for the rest of their careers. Buyers who punish it get optimism, and optimism is expensive at build scale.

Bottom line

You are not being asked to estimate a project. You are being asked to estimate a project whose difficulty is unknown, which is a different task with a different method. Split the unknown into its four parts, run the cheap decisive test for each — three of which need no vendor at all — find the ceiling from your own experts before you set any target, fix the time and money rather than the scope, and write the stopping rule while everybody is still calm. Do that and the estimate arrives with real numbers behind it in six weeks. Skip it and you will get an estimate in six days, and it will be a guess wearing a spreadsheet.

Frequently asked questions

How much should a scoping or feasibility phase cost?

As a rule of thumb, ten to twenty percent of the projected build, over four to six weeks. The test is not the percentage though — it is whether the deliverable is a decision. If the output is a written finding with measured numbers, a cost per item and a recommendation the firm will stand behind, it is scoping. If the output is a partly built system, it is the build with a different name and it should be priced and governed as one.

What accuracy should I ask for?

Derive it rather than choose it. Establish the ceiling by measuring how often two of your own experts agree on the same cases, measure what your current process achieves, then set a threshold from the cost of a wrong answer versus the cost of a missed one. Those two costs are rarely equal, and once they are written down the target usually stops being a single number and becomes an operating point with two figures.

Can a vendor give a fixed price before scoping?

They can, and the price will carry their uncertainty as a risk premium, or it will be low and become a change-order conversation. Neither is dishonest; both are expensive. The better shape is a fixed price for the scoping work, where the unknowns are bounded by time rather than by scope, followed by a fixed price for the build once the numbers exist. Firms are usually happy to commit hard to the second once they have done the first.

How small can the first slice be and still be useful?

Small in breadth, never in difficulty. One document type, one region, one product line is fine. Clean data, curated examples and the easy cases removed is not, because the hard cases are the reason the project exists. The test of a slice is whether a real decision follows from the result and whether somebody actually benefits if it works.

What should I do myself before hiring anyone?

Three things, and they take days rather than weeks. Pull real records through the real approval path, so data access is proven rather than assumed. Have two experienced people label the same hundred cases independently and measure their agreement. Measure the current process error rate on those same cases. Arriving at a vendor conversation with those three numbers changes the quality of every proposal you receive.

1 business day response

Trying to size something nobody can size yet?

Tell us what happens today and what you wish were different. We will name the unknown that decides the project, the cheapest test that resolves it, and which parts of that test you can run yourselves before hiring anyone.

Email an engineerHow we workMore insights →
FeasibilityMachine LearningData EngineeringDelivery