Skip to main content
Commercial

When a proof of concept should be killed

Most concepts are not stopped and not finished. They are extended. The sponsor who has to decide whether to fund the next increment is answering a narrower question than the room usually asks, and answering it well takes about one page, one meeting, and a set of thresholds that were written down before anybody had a reason to argue about them.

The question a sponsor is actually answering

A continue-or-kill review usually opens with the wrong question. Somebody asks whether the concept is working, and the room spends ninety minutes on a word nobody defined. The question a sponsor is paid to answer is narrower and much easier to settle: if this proposal arrived on my desk today, priced at the money that is still unspent, with everything I now know about the data, the workflow, and the team, would I approve it? That reframing does the entire job. It removes the money already spent from the calculation, because that money is gone under every option. It replaces a verdict on the past with a decision about the future, which is the only thing a budget can change.

Sponsors resist the reframing because it feels like it discards information. It does not. Everything learned during the concept is still in the room, and it is worth more now than it was at kickoff, because the estimate of feasibility is grounded in the organization's own data rather than a vendor's reference case. What gets discarded is only the emotional weight of the invoices, and that weight is the single most reliable source of bad capital allocation in technology programs.

The rest of this piece is the operating detail: where a concept can fail and what each kind of failure implies, how to write kill criteria that survive contact with a sympathetic room, which signals mislead, what to salvage on the way out, and the contract mechanics that let a sponsor stop without a fight. It is written for the person holding the budget, and it is equally useful to the partner or vendor on the other side of the table, because a supplier who can name the kill criteria out loud is a supplier who gets funded again.

A concept can fail in four places, and only one of them is fatal

"It did not work" collapses four separate findings into one sentence, and the four have completely different consequences. Feasibility failure means the technique cannot do the task at the quality the mission requires. Value failure means it can, but the gain is not worth what it costs. Integration failure means it works in the notebook and cannot reach the system of record. Operability failure means it works and cannot be run by the people who would have to run it every day.

Only feasibility failure is a genuine stop. The other three are pricing and sequencing problems, and they are frequently solvable by a different, smaller engagement rather than an extension of the current one. A sponsor who cannot say which of the four is in front of them is not ready to decide, and the correct move in that situation is a two-week diagnostic, not a six-month extension.

Where it failedWhat the evidence looks likeWhat it implies
FeasibilityQuality plateaus well below the threshold across several honest attempts, and the error pattern is intrinsic to the task rather than to a fixable pipeline defectStop. Extensions buy variance, not capability
ValueThe threshold is met, and the modeled annual benefit does not clear the modeled annual cost of running it, including review laborStop this scope. Re-aim at a use case with a larger denominator
IntegrationResults are real on exported data; the write path into the system of record does not exist, or the data owner has not agreed to itNot a technology decision. Fund an integration increment or stop for organizational reasons
OperabilityIt works when the vendor's engineer runs it; nobody on staff can retrain it, monitor it, or explain an output to an auditorPrice the operating model honestly, then re-decide. Often smaller than it looks
DefinitionNobody can produce the written claim the concept was supposed to test, or the metric changed during the engagementStop and restart with a written claim. Do not extend an undefined experiment
Data accessThe representative data never arrived, and the concept ran on a sample the owner will not vouch forThe blocker is a permission, not a model. Solve it or shut down

Write the kill criteria before the first invoice

Kill criteria are worthless when they are written by a room that already knows the result. Written at kickoff, when nobody's reputation is attached to an outcome, they are the cheapest governance a program will ever buy. Three items make them real, and all three have to be on the same page: the threshold, the date, and the name of the person who decides.

The threshold. A number, on a named metric, measured on data the buyer selected. "Meaningful improvement" is not a threshold. "At least 20 percentage points above the current process on the 400-item review set, scored by the same rubric the team uses today" is. Write the comparison baseline into the same sentence, because a concept with no measured baseline cannot produce a result that clears or fails anything.

The date. A calendar date for the decision, chosen so the decision lands before the next budget commitment rather than after it. Concepts drift because their end is defined by an event that keeps moving. A date does not move.

The decider. One named person with the authority to stop the spend, and a stated rule for what happens if that person is unavailable. Programs where the decision authority is "the steering committee" continue by default, because a committee that cannot reach consensus has continued.

The technique that makes people write honest thresholds is the premortem, described by Gary Klein in Harvard Business Review in September 2007: at kickoff, tell the team to assume the effort has already failed, then have each person write down why. It surfaces the reservations that politeness suppresses at kickoff and that hindsight makes unarguable at the gate. Run it for twenty minutes, keep the list, and read it aloud at the review. Half of the reasons a concept eventually stalls are usually on that page, written by the people now arguing to continue.

The gate itself is an old idea with a good record. Robert G. Cooper described Stage-Gate in Business Horizons in 1990 as a structure where work proceeds in stages separated by decision points with explicit go/kill criteria. The structure is not the hard part. The hard part is that a gate with no kill option is a status meeting wearing a gate's clothes, and organizations quietly convert one into the other over about three cycles.

What should carry weight at a continue-or-kill gate

Measured result against the threshold written at kickoff
96%
Whether the data the concept needs actually exists and is accessible
91%
Annual run cost against annual benefit, both modeled
87%
A named owner willing to run it after handover
80%
Write path into the system of record, agreed by its owner
76%
Team credibility and responsiveness during the engagement
64%

Editorial weighting of how much each item should influence the gate decision. Illustrative ranking, not a measured statistic.

The ranking above is a claim about decision weight, not a measurement of anything. It says that the written threshold outranks every other input at the gate, that data access and run economics outrank organizational comfort, and that team quality belongs last. Team quality belongs last not because it does not matter, but because it is the input most likely to keep a doomed effort alive. Good engineers make failing work feel like progress, and a sponsor who weights that highly will fund the same concept three times.

What the record says about why concepts fail

The most useful published study on this is the RAND report The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (RR-A2680-1, August 2024), by James Ryseff, Brandon F. De Bruhl, and Sydne J. Newberry, built from interviews with 65 data scientists and engineers who each had at least five years of experience building models in industry or academia. The five root causes it identifies in industry are worth reading against any concept currently under review.

First, leadership fails to communicate what problem is being solved and which metric matters, so teams optimize the wrong thing or build something that does not fit the workflow. Second, the organization lacks enough quality data to train an effective model. Third, teams focus on the newest technology rather than on the user's problem. Fourth, infrastructure for managing data and deploying models is underfunded. Fifth, the problem is simply beyond what the technology can currently do.

Note what is on that list and what is not. Four of the five causes are decisions the buying organization made before any model was trained. Only the fifth is a property of the technology. A sponsor reviewing a stalled concept should therefore spend most of the meeting on their own side of the table, and the honest version of many kill decisions is that the concept was never given a defined problem or usable data.

RAND also offers a rule that doubles as a kill criterion at kickoff: choose enduring problems, and be prepared to commit a product team to a specific problem for at least a year, because a project not worth that commitment most likely is not worth starting. Read backwards, that says a concept whose sponsor would not fund a year of work on the same problem should be stopped now regardless of how the numbers came out.

One caution about statistics in this space. Widely repeated figures on technology project failure rates are contested, and the sources behind them frequently use inconsistent definitions of failure and unrepresentative samples. The RAND report itself introduces the "more than 80 percent" figure as an estimate attributed to others rather than as its own measurement. Treat any single headline percentage as an argument, not evidence, and make the decision on the organization's own measured result instead.

Six signals that mean stop

The quality curve is flat across honest attempts. Three or four serious approaches, each with a real change in method rather than a parameter tweak, all landing in the same band well below the threshold. Flatness across genuinely different attempts is the strongest evidence a concept produces.

The representative data does not exist. Not "is hard to get" but does not exist in the volume, quality, or labeling the task requires, and creating it would cost more than the benefit. This is a finding, and it is often the single most valuable output of a concept.

The metric moved after the results came in. If the team is proposing a new measure of success at the gate, the original measure was not met. That may be defensible, and it is a restart with a new written claim, never an extension of the current engagement.

The run cost exceeds the benefit at realistic volume. Concepts are priced at concept volume. Multiply inference, storage, monitoring, and the human review the workflow will actually require by production volume, then compare against the labor the system would displace. Many technically successful concepts die correctly here.

The workflow owner has not agreed to change the workflow. A concept whose value depends on people doing their jobs differently, where the manager of those people has not committed to the change, has an unfunded dependency larger than the engineering. That is an organizational decision, and it should be made before more engineering money moves.

The champion has left. Unsentimental and consistently true. When the executive who wanted it moves on and no successor claims it, the concept has lost the thing that would have carried it through the friction of deployment. Stop it and let the next owner start something they chose.

A gate with no kill option is a status meeting wearing a gate's clothes, and organizations quietly convert one into the other over about three cycles.

Four signals that look like failure and are not

The first numbers are bad. Early results on a pipeline that has not been debugged tell you about the pipeline. The question is whether the errors are structural or mechanical. Mechanical errors have identifiable causes with named fixes and estimated effort. Structural errors are described in adjectives.

It works on the easy half. A concept that handles 60 percent of cases well and fails on the messy remainder has often found the shape of the real product, which is a system that routes the easy cases and escalates the rest. Ask what the value is of automating only the clean 60 percent. Frequently that is the whole business case, and the attempt at full coverage was the design error.

Users disliked the interface. Interface complaints during a concept are a design finding, not a capability finding, and they are the cheapest problem on this page to fix. Separate them from output quality before they contaminate the verdict, because they arrive loudly and land early.

It is late. Schedule slip on a concept usually means data access took longer than planned, which is a fact about the organization. Late plus on-threshold is a success with a scheduling lesson. Late plus below-threshold is a failure that the lateness is being used to explain away.

The tells that a concept is alive for reasons other than evidence

Escalation of commitment is well documented in the decision literature. Hal R. Arkes and Catherine Blumer described the pattern in "The psychology of sunk cost" in Organizational Behavior and Human Decision Processes in 1985: people continue an endeavor once money, effort, or time has been invested, and the prior investment changes the choice even though it should not. Forty years later it remains the most reliable prediction anyone can make about a review meeting, and it shows up in a small number of recognizable forms.

The goal is now further away than at kickoff. The concept was to prove a claim in ninety days. At day ninety the plan is a six-month phase to prove a broader claim. Scope that grows at the gate is scope that is buying time.

The demo set is curated. Every review shows outputs the presenter selected. A concept that cannot be run live on items the sponsor picks in the room has not been evaluated; it has been exhibited.

The success language has gone qualitative. Kickoff talked about accuracy on a review set. The gate talks about learnings, capability building, and positioning. Those may be real, and none of them was what the budget bought.

Nobody in the room is neutral. The vendor is paid to continue, the internal team is staffed on it, the champion sponsored it, and the analyst wrote the deck. If nobody present would be fine either way, invite somebody who would be. An outside read costs a fraction of one extension and is the only input in the room with no position.

The stopping cost is unknown. When nobody can say what stopping would cost, continuation wins by default. Price the stop before the meeting: termination terms, remaining commitments, what the team does next week. A stop with a known price competes fairly; an unpriced one never does.

Harvest before shutdown

A concept that is stopped well returns most of its cost in assets that outlive it. A concept that is stopped badly returns a folder of dead links. The difference is about a week of deliberate effort, scheduled before the shutdown rather than after, while the people who built it are still under contract and still remember why things were done the way they were.

  • The measured baseline. What the current process scores on the same items under the same rubric. Organizations almost never have this, it does not expire, and it makes the next vendor evaluation cheaper and faster.
  • The labeled data. Whatever was labeled or curated during the engagement, with the labeling instructions and the disagreement rate between labelers. This is usually the single most expensive artifact produced.
  • The evaluation harness. The scripts and held-out set that produced the numbers, runnable by someone who was not on the team. A reusable harness turns every future claim into a testable one.
  • The data map. Where the records live, who owns them, which fields are trustworthy, what the access path was, and how long each approval took. This is institutional knowledge that normally leaves with the contractor.
  • The negative result, in writing. One page: the claim, the method, the measured result, and the specific reason it did not clear. Without it the organization will fund the same idea in eighteen months at full price.
  • The cost model. Actual compute, storage, and human review time observed during the concept, extrapolated to production volume, with the assumptions listed.
  • Access and credential closeout. Revoke keys, remove accounts, confirm deletion or return of any copied data under the terms of the agreement, and get written confirmation.
  • The disposition record. A dated note of the decision, the criteria applied, who decided, and what happens to the artifacts. This is what makes the next review faster.

For anyone whose organization already works to a recognized risk framework, the discipline has a home in it. The NIST AI Risk Management Framework (AI RMF 1.0, NIST AI 100-1, January 2023) states under MANAGE 2.4 that mechanisms are in place and applied, and responsibilities assigned and understood, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use. MANAGE 4.1 names decommissioning among the post-deployment mechanisms an organization is expected to have. The ability to turn something off is treated there as a control, not a failure, which is the correct posture and a useful thing to be able to cite in a room that treats stopping as an admission.

The contract mechanics of stopping

The cleanest kill decisions are made possible months earlier, in how the work was bought. The lever is increment size. Money committed in short, separately priced increments can be stopped at a boundary; money committed as a single multi-quarter scope can only be stopped by a negotiation. Sponsors who intend to keep the option to stop should buy in increments that each produce something usable on their own.

Federal buyers have this written into the regulation. FAR 39.102 directs agencies to analyze risks, benefits, and costs before contracting for information technology, and names prototyping prior to implementation, modular contracting, and post-implementation reviews among the techniques for managing that risk. FAR 39.103 implements 41 U.S.C. 2308 and requires that each increment comprise a system or solution that is not dependent on any subsequent increment in order to perform its principal functions. That sentence is the whole design principle: an increment that only has value if the next one is funded has removed the option to stop.

The statute goes further than most people expect. Under 41 U.S.C. 2308(c)(2), a contract for an increment should be awarded within 180 days after the solicitation is issued, and if it cannot be, the increment should be considered for cancellation. Cancellation is not an exception in the modular contracting scheme. It is the named response to an increment that has stopped moving.

InstrumentHow the sponsor stopsWhat it typically costs to stop
Commercial fixed-price, milestone-billedDo not authorize the next milestone; the engagement ends at the boundaryWork completed through the last accepted milestone. The cleanest structure available
Commercial time and materialsStop-work notice under the master agreement, subject to its notice periodHours worked plus the notice period. Negotiate a short notice window at signature
Federal commercial acquisition, FAR 52.212-4Termination for the Government's convenience under paragraph (l)A percentage of the contract price reflecting work performed before the notice, plus reasonable charges the contractor can demonstrate from its standard records
Federal fixed-price, FAR 52.249-2Contracting officer issues a Notice of Termination specifying extent and effective date, when termination is in the Government's interestA settlement process under FAR part 49. Slower and more procedural than the commercial clause
Modular IT increments, FAR 39.103Decline to award the next increment; each increment stands on its ownNothing beyond the increment already awarded, which is the point of the structure
DoD prototype transaction, 10 U.S.C. 4022Do not exercise the follow-on production path; terms of the agreement govern the prototype effort itselfSet by the agreement, not by the FAR termination clauses

The prototype line deserves a note, because it is where the incentive to declare victory is strongest. Under 10 U.S.C. 4022(f)(2), a follow-on production contract or transaction may be awarded to the participants without competitive procedures only if competitive procedures were used to select the parties for the prototype transaction and the participants successfully completed the prototype project. Subsection (f)(3) leaves the determination of successful completion with the Department. That structure puts real money behind the word "successful," which is exactly why the completion criteria belong in the agreement at the start rather than in a memo at the end.

A ninety-day concept with gates that can actually close

The shape below is a working default for a concept that is meant to end in a decision. The durations move with the domain; the gates do not. What makes it function is that each gate has a criterion that was written before the gate arrived, and that stopping at any of them is a normal outcome rather than an escalation.

Gate structure for a decision-grade concept

1
Written claim, threshold, decision date, named decider, premortem list
Before any spend
2
Data gate: representative records in hand, owner's approval documented
Day 10 to 15
3
Baseline gate: the current process measured on the same items
Day 20 to 25
4
Signal gate: is the quality curve rising or flat across distinct attempts
Day 45
5
Scored run on buyer-held items, frozen version, rubric fixed in advance
Day 75 to 85
6
Decision: continue, restructure, or stop and harvest
Day 90

Gate 2 is the one most often skipped and the one that saves the most money. A concept that reaches day 15 without representative data has already discovered its answer, and every dollar after that is spent on a question the organization is not able to ask yet. Stopping at gate 2 costs about a sixth of the engagement and produces a finding the organization needs regardless of which vendor it eventually chooses.

What a partner should read into a killed concept

Sponsors underrate how much information a kill decision broadcasts to the supplier market, and suppliers underrate how much a clean kill helps them. A sponsor who stops on written criteria, pays through the boundary promptly, and writes down the finding is a sponsor that good firms will bid on again. A sponsor who extends indefinitely, then stops abruptly and disputes the last invoice, gets priced accordingly on the next engagement and loses access to the firms that have alternatives.

For the supplier, the useful posture is to bring the kill criteria up first. Naming the conditions under which the work should stop, in the proposal, reads as confidence rather than pessimism, and it converts the sponsor's largest unspoken fear into a term of the agreement. A firm that has a written negative result from a prior engagement has something rarer than a case study: evidence that it will report a result the client did not want. That is the quality a sponsor is buying when the stakes are real.

Both sides should also separate the concept's verdict from the relationship. A concept that ends at gate 2 because the data was not accessible has told everyone something true and nothing about the team's competence. The right follow-up is often a different, smaller engagement aimed at the blocker, run by the same people, who now know the organization's systems better than anyone else available.

Bottom line

Decide in advance and the decision is arithmetic. The threshold, the date, and the decider go on one page at kickoff; the review answers whether the same money would be committed today with today's information; the four kinds of failure get separated, because only one of them is a stop; the assets get harvested while the team is still under contract; and the contract was structured in increments so that stopping is a boundary rather than a negotiation. What makes this hard is never the analysis. It is that continuing feels like commitment and stopping feels like admitting error, when in fact stopping a concept that has produced its answer is the cheapest correct decision a sponsor makes all year.

Frequently asked questions

How long should a proof of concept run before a decision?

Long enough to get representative data, measure the current process, and run one scored evaluation, which is commonly 60 to 120 days for a narrow claim. The duration matters less than the decision date being fixed at kickoff. Set intermediate gates at data access and at baseline measurement, because a concept that cannot clear those has already produced its answer at a fraction of the cost.

What is the difference between killing a concept and pausing it?

A pause keeps the team, the access, and the budget line reserved while nothing is being learned, which is the most expensive state a concept can be in. If the blocker has a named owner and a date, that is a schedule change. If it does not, it is a kill being described gently. Stop it, harvest the artifacts, and restart later with a fresh decision when the blocker clears.

Who should decide whether to continue funding a proof of concept?

One named person with authority over the budget line, identified in writing before the work begins, with a stated alternate. Committees continue by default, because failing to reach consensus is indistinguishable from a decision to keep going. The vendor, the internal delivery team, and the original champion all have positions and should inform the decision rather than make it.

What does it cost a federal sponsor to stop a concept mid-engagement?

It depends on the instrument. Under FAR 52.212-4(l), for commercial products and services, the contractor is paid a percentage of the contract price reflecting the work performed before the notice of termination plus reasonable charges it can demonstrate from its standard records. FAR 52.249-2 governs fixed-price terminations for convenience through a settlement process under FAR part 49. Modular increments under FAR 39.103 are the cheapest to stop, because declining the next increment costs nothing beyond what was already awarded.

Is a failed proof of concept a waste of money?

Only if nothing is kept. A concept that ends in a stop should leave behind a measured baseline for the current process, the labeled data, a runnable evaluation harness, a map of where the data lives and who owns it, a modeled run cost, and a one-page written record of the claim and why it did not clear. Those assets make every subsequent evaluation cheaper, and the written negative result is what stops the organization from buying the same idea again in eighteen months.

1 business day response

Deciding whether to fund the next increment?

We build AI, ML, data, and cloud systems for federal and commercial missions, and we run independent reads on concepts that have stalled: what was measured, what was not, and what the remaining money would buy.

How we workMore insights →Start a conversation
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE