The invoice is the smallest number in the file
A proof of concept gets approved on one number. Six to eight weeks, a fixed fee, some model and infrastructure spend on top. That number is defensible, it fits inside a director's signing authority, and it is almost never the number that decides whether the work was worth doing. The costs that decide that show up on the platform team's quarter, in a security questionnaire eleven months later, and in the meeting where somebody asks what happened to the last two of these. None of them appear in the approval.

This is not an argument against running one. A proof of concept is the correct instrument when a specific decision is waiting on a measurable unknown, and we build them. It is an argument for pricing the whole thing, because the mispricing is systematic and it runs in one direction. Every cost that lands outside the engagement is invisible at approval time and unavoidable afterward, so organizations reliably run more of these than they can absorb and get less out of each one than they paid for.
The ledger below is seven lines. Each one is concrete, each one has a way to price it in advance, and the pricing takes a few hours of writing rather than a governance process. The pattern of who pays is the useful part: in a healthy engagement roughly a third to a half of the real cost sits outside the contract entirely.
You are probably here because
- The invoice came in on budget and the quarter still slipped, because the people who know the data were never actually released to work on it.
- The readout said the system was right eighty-seven percent of the time, and nobody could say what the current process gets right.
- Someone asked what happened with the last two, and the room had no answer.
- An access review turned up a service account and a copy of production data from an experiment that ended last year.
The seven cost lines below take these one at a time and give each one a way to price it before the work starts. They share a root cause: the work was scoped as something to build rather than as a decision to settle.
Where the true cost of a six-week proof of concept lands
Planning split we use for sizing. Loaded internal rates, six to eight weeks of working time. Shares move with data readiness, not with model choice.
Cost one: the people you cannot backfill
The team running the proof of concept is not the expensive part. The expensive part is the four or five people inside the company it has to borrow. There is always a staff engineer who is the only person who knows why the orders table has two customer identifier columns and which one is authoritative. There is a domain expert whose judgment is the ground truth, because whatever the model produces has to be compared against what a competent human would have decided. There is a platform engineer who has to cut a service account, open a network path, and answer for it. There is usually an analytics engineer who owns the dbt models everything reads.
The typical draw is four to ten hours a week from each of them for the duration, plus two concentrated blocks: a few days at the front for data access and schema explanation, and a few days at the end for review and adjudication of disagreements. Priced at loaded internal rates that lands between a quarter and a half of the external fee. It does not come out of a discretionary line. It comes out of the roadmap, and it comes out of the specific people who were already the constraint on the roadmap.
That last point is what makes this line worse than it looks in dollars. Every company's delivery schedule is set by a small number of individuals, and the proof of concept selects for exactly those individuals, because they are the ones who know how the data works. Four hours a week of a staff engineer is not four hours of throughput. It is four hours plus the context-switching around them, on the person whose queue everything else is sitting in.
Price it by naming them. Put the internal contributors in the engagement plan by name with a weekly hour figure next to each, and have the manager whose roadmap they sit on sign that figure. If no manager will sign the hours, the work has been approved but not funded, and it will run three weeks late for reasons that get attributed to the vendor.
Cost two: the access you opened and never closed
To get data in front of the system, somebody made an exception. The shapes repeat across every company we have worked in. A copy of production data lands in a lower environment because masking it properly was a two-week project nobody had. A service account gets read on a whole schema because scoping it to eleven tables required a conversation with three teams. A storage bucket or a warehouse share gets a policy written to unblock the week. An extract sits on a laptop. An API key goes in a dotenv file and then into a message thread so a second engineer can run it.
The engagement ends. The exception does not. Nobody revokes it because revoking it is nobody's job, it breaks nothing when it stays, and no test fails. It surfaces later in an access review, a SOC 2 audit, or a customer security questionnaire that asks where their data has been copied. The remediation itself is an afternoon of work. The cost is the finding, the evidence you now have to produce, and the conversation with a customer whose records were in a sandbox copy for a year because of an experiment that was cancelled. The exception you made to get the data out is the part of the proof of concept that outlives it.
Fixing this is cheap and almost nobody does it. Every access grant made for the work gets a ticket with an expiry date, created at the moment the grant is made, not at the end when everyone has moved on. The final milestone of the engagement is not the readout. It is access revoked, extracts destroyed, evidence attached, signed by whoever owns the systems. Bill it as a real day of work so it survives a schedule squeeze, because a close-out that is treated as a courtesy is the first thing dropped when the readout slips.
A proof-of-concept copy is a processing activity, not a scratch file
If the data includes personal information, the copy has a purpose and a retention period whether or not anyone wrote them down. GDPR Article 5(1)(b) binds the copy to the purpose it was collected for, Article 5(1)(e) requires it not be kept longer than necessary, and Article 30 expects the activity to be in your records of processing. There is no experimentation carve-out. Under a SOC 2 engagement the same copy becomes an access-control and data-lifecycle question at the next audit window, and under HIPAA a vendor touching protected health information needs a business associate agreement in place first, not retroactively. The practical version: decide the deletion date before the extract is made, and put it in the ticket.
Cost three: the decision that stopped moving
A proof of concept is supposed to inform a decision. In practice, starting one suspends it. Nothing in the organization moves on a question that is officially in flight, because no manager wants to make a call that a running evaluation might contradict next month. That freeze is a real cost and it is longer than the engagement.
Count the calendar honestly. Three weeks pass before kickoff while data access clears legal and security. Six weeks of work. Then the readout waits for a monthly forum, and the decision waits on the forum after that, and if money is attached it waits for a planning cycle. A six-week engagement routinely freezes a decision for four to five months. During that window the teams downstream do not wait. The manual process acquires a spreadsheet that someone now maintains, and a neighboring team buys a point tool on a card because their problem was urgent. You come out of the evaluation deciding against a baseline that moved, with two workarounds to unwind that nobody counted as a cost of the evaluation.
Price it by writing the decision date before the start date. Name the person who makes the call, name the forum, and put the date in the plan. If that date is more than a month after the readout, the freeze is longer than the work, and the honest move is a shorter and coarser evaluation that lands inside the window the organization can actually act in. A rough answer you can act on in March beats a precise one you act on in July.
Cost four: credibility, and why the third one is the hard one
Every organization has a finite number of times it will fund this class of work before the phrase itself starts costing you. The first proof of concept gets curiosity. The second gets scrutiny. The third gets the question that ends the meeting: what happened with the previous two? That question is usually unanswerable, and not because the previous two failed. It is unanswerable because nobody wrote down what they were supposed to establish, so there is no record of what was learned, only a memory of a demo that looked good.
The mechanism is worth being precise about. A proof of concept that ends in a demonstration rather than a number cannot be defended a year later by anyone, including the people who ran it. The demo was genuinely impressive, everyone remembers that part, and nobody can state what it proved. So the next funding request inherits an unpaid debt: it has to re-argue a case that was never settled, in front of a group that now associates the work with expense and no decision.
The insurance is one sentence, written before any work starts, in a fixed form. We will proceed if the system reaches a stated level on a stated metric, measured on a holdout the buyer controls, and we will stop if it does not. Then publish the answer either way, in a short memo, whichever direction it lands. A published negative result costs one proof of concept. An impressive demo that settled nothing costs the next two.
Cost five: the prototype that quietly becomes the product
The thing works. Somebody shows it in a leadership review, or worse, to a customer. Now there is pressure to put it in front of a handful of real users, just to see. This is the most expensive line on the ledger and it is entirely avoidable, because it is a decision that gets made by drift rather than by anyone choosing it.
What actually exists at that moment is a Streamlit app with a key in the environment, a notebook that runs correctly on one laptop, a vector index that was built once by hand with no rebuild path, prompts living in Python string literals with no version history, and an evaluation that is three examples in a docstring. None of that is a criticism. It is the correct shape for something whose job was to answer a question in six weeks. It is a bad shape for something a person depends on to do their job.
The cost is not rewriting it. A small system is cheap to rewrite, and a good proof of concept is small. The cost is what accumulates between the demo and the rewrite. Users arrive, and users generate expectations. Someone wires a second system to its output. It starts writing data that becomes authoritative somewhere because a report reads it. Once three things depend on it you are not rewriting, you are migrating, and migration is a different order of work with a coordination cost that has nothing to do with the code.
Four questions separate a throwaway you can actually throw away from one you cannot. Does it write data that anything else reads? Does anyone outside the team have a URL for it? Is it wired to your identity provider, meaning someone treated it as a real application? Does any scheduled job depend on its output? If all four are no, deleting it costs nothing. The moment any is yes, you own it, and the honest response is to stop calling it a proof of concept and fund it as a product.
Cost six: the baseline you skipped
This one gets paid at the readout, and it is the one that most often makes the whole exercise worthless. A proof of concept that reports the system was right eighty-seven percent of the time has reported nothing, because there is no second number. What is the current process right? Almost nobody knows. That figure exists as folklore, not as a measurement, and the folklore is usually generous.
Measuring it is a week of unglamorous work: sample a few hundred real cases from the current process, have two qualified humans adjudicate each independently, resolve the disagreements, and count. You get the accuracy, and you also get the inter-rater agreement, which is the more useful number, because it tells you the ceiling. If two experienced people disagree on eighteen percent of cases, no system is going to be measured above the low eighties on that task and any vendor promising ninety-five is measuring something else.
A skipped baseline is not deferred. Once the team has read the model's output, the comparison it was supposed to make no longer exists. A human reviewing a suggested answer agrees with it far more often than a human deciding from scratch, so the adjudication is contaminated the moment anyone has seen a prediction. The process itself has usually drifted as well, since people work more carefully while a project is watching them. The measurement you postponed is not waiting for you at the end. It is gone.
The same logic applies to the holdout. Pull a set aside before anything is built, keep it out of the working environment, and evaluate against it once, on a frozen version, at the end. Every extra look is a tuning signal, and after three or four looks the holdout is training data wearing a different name. This costs an hour of discipline in week one and is the difference between a result and a story.
Send it over and we will tell you what we would change.
Email the scope or statement of work for the proof of concept you are about to approve, plus the one decision it is supposed to settle, to contact@precisionfederal.com. You get back a short written note naming the three things we would change and why. One business day. No charge, no meeting, no deck.
contact@precisionfederal.comCost seven: the answer has a shelf life
A result is a measurement of one configuration, on one dataset, at one date. Three things move underneath it. Model versions get retired on published schedules and the replacement does not behave identically, so prompts and thresholds tuned against one version need revalidation against the next. The data moves, because schemas change, a product launch shifts the distribution, and seasonality does what it does. The process moves, because the people doing the work change how they do it, sometimes in response to the project itself.
The practical rule we use: a result older than about two quarters is a hypothesis again, not evidence. If a decision is made nine months after the evaluation on the strength of that number, it is being made on a number nobody has checked, and the checking is now expensive because the environment has to be rebuilt.
What does not decay is the apparatus. The evaluation harness, the labeled set, the adjudication guidance that produced the labels, and the measured baseline stay valuable for years and get reused by every subsequent effort against that problem. That is the strongest argument for where the money should go. Spend the engagement on the harness and the labels rather than on squeezing three more points out of a model, because the three points expire and the harness does not.
The ledger, priced up front
| Cost line | Whose budget it lands on | When it shows up | How to price it in advance |
|---|---|---|---|
| Internal senior attention | The roadmap, not the project | During, and for a quarter after | Named people, weekly hours, signed by their manager |
| Access left open | Security and compliance | Next access review, audit, or customer questionnaire | Expiry ticket at grant time; close-out as a billed milestone |
| Frozen decision | Every team downstream of the call | From kickoff until someone actually decides | Set the decision date and the decision owner before the start date |
| Spent credibility | The next initiative | At the funding request after this one | Write the proceed-or-stop sentence first; publish the answer either way |
| Prototype debt | Engineering, later and larger | The first time a real user depends on it | Choose the shape up front; keep the throwaway unwired and dated |
| Missing baseline | The engagement itself | At the readout, when the number means nothing | One week of measurement in week one, before anything is built |
| Result decay | Whoever inherits the decision | Six to nine months out | Fund the harness and labels, not the last three points of accuracy |
Three shapes, and only one of them is a proof of concept
Most of the damage above comes from one ambiguity: the phrase covers three different things with different costs, and the sponsor and the engineers frequently mean different ones. Naming the shape in the first meeting removes more waste than any process change.
| Shape | What it establishes | Calendar and cost driver | What it leaves behind |
|---|---|---|---|
| Demonstration An existence proof | That the thing can be made to work once, on inputs you selected | One to two weeks, driven only by engineer time | Screenshots and a recording. Nothing reusable, and that is fine if everyone knows it |
| Measured evaluation The actual proof of concept | That it beats the current process by a stated margin on data it has not seen | Six to eight weeks, driven by data access and label adjudication | Harness, labeled set, measured baseline, a written answer that survives either result |
| Thin production slice | That it survives real inputs, real latency, real failure, and a real user | Eight to twelve weeks, driven by integration and review cycles | Running code, monitoring, an on-call runbook, one narrow path in production |
The failure mode for each is specific. A demonstration goes wrong when it reaches a customer and becomes a commitment. A measured evaluation goes wrong when there is no baseline, which turns it back into a demonstration with more steps. A thin slice goes wrong when it is scoped as one path and built like a platform, which is how a twelve-week commitment becomes a nine-month one.
What actually decides whether it pays back
Across the engagements we have run and the ones we have been brought in to clean up afterward, the variables that separate a proof of concept that produced a decision from one that produced a memory are not technical. Model choice barely registers. These are the ones that do, weighted the way we weight them when sizing the work.
What determines the return on a proof of concept
Weights sum to 100. Every row is decided in the first week, before a line of the system exists.
Where thirty working days actually go
Sponsors budget the calendar as though it is mostly building. It is mostly not. Here is the split we plan against for a six-week evaluation, ordered by size rather than sequence, and it is the reason a schedule that assumes four weeks of model work slips before anyone writes code.
Thirty working days, by activity
Ordered by size, not sequence. Access and data preparation are half the calendar in most engagements.
An engagement with the quiet costs priced in
Six weeks, plus the two that nobody schedules
Write these down before the first week
- The decision this informs, the person who makes it, and the date it gets made
- The sentence that ends it: we proceed if the system reaches X on metric Y, measured on holdout Z
- The measured baseline for the current process, and who is sampling and adjudicating it
- Named internal contributors with weekly hours, signed by the manager whose roadmap they sit on
- Every access grant paired with an expiry date and a ticket, created when the grant is made
- Which of the three shapes you are building, and what gets deleted at the end
- The artifacts that survive either answer: harness, labeled set, baseline, adjudication guidance, memo
How organizations pay these costs without noticing
- Funding the engagement and not the internal hours. The vendor is paid and the schedule still slips, because the people who know the data were never released to work on it.
- Starting the clock before the data path is approved. Three of the six weeks go to a legal review that could have run in parallel a month earlier.
- Measuring the baseline after the model exists. The adjudication is contaminated by the model's suggestions and the comparison is no longer meaningful.
- Evaluating repeatedly on the holdout. After the fourth look it is training data, and the reported number is optimistic by an unknown amount.
- Letting the demonstration reach a customer. A commitment is created by a screen share, and the shape of the work changes without anyone deciding to change it.
- Leaving the sandbox copy in place because it might be needed again. It is never needed again and it is always found by an auditor.
- Ending at a demonstration instead of a number. Nothing is settled, so the next request has to re-argue the same case with less goodwill.
- Treating a nine-month-old result as current. The model version, the data distribution, and the process have all moved since it was measured.
Bottom line
The engagement fee is the part of a proof of concept that is easiest to see and least likely to determine the outcome. The rest of the ledger is senior attention taken off the roadmap, access nobody closes, a decision that stops moving the day the work starts, credibility you can only spend a few times, a prototype that turns into a product by drift, a baseline that cannot be recovered once it is skipped, and an answer that expires. All seven can be priced in a few hours of writing before anyone opens an editor. Do that and the number you approve is close to the number you pay, which is the whole point of approving it.
Frequently asked questions
Budget the external fee, then add roughly the same again for everything outside the contract. In our sizing the invoice is a bit under forty percent of the true cost, with internal senior engineering and domain time the next largest line. If a plan shows the vendor fee as the whole cost, the plan is missing about half the number.
Six to eight weeks of working time is right for a measured evaluation, but the length that matters is kickoff to decision. Add the access lead time in front and the decision latency behind, and a six-week engagement often occupies four or five months of organizational calendar. Set the decision date first and work backward.
The evaluation harness and its scoring script, the labeled set with the adjudication guidance that produced it, the measured baseline for the current process, the error analysis, and a short memo stating what was established. Plus evidence that access was revoked and extracts destroyed. Those artifacts get reused by the next attempt at the same problem, which is why a negative result still returns something.
Decide the shape in the first meeting and keep the throwaway genuinely disposable: no writes that anything else reads, no URL outside the team, no identity-provider integration, no scheduled job depending on its output, and a deletion date in the plan. The moment any of those changes, stop calling it a proof of concept and fund it as a product with tests, monitoring, and an owner.
When the decision would not change either way, when nobody will commit the internal hours, or when the decision cannot be made within a month of the readout. In the last case the freeze costs more than the evidence is worth, and a shorter coarser evaluation that lands inside the window you can act in is the better trade.
