Skip to main content
Business Development

What an AI proof-of-value should cost and how long it should take

Buyers get quotes for a first AI engagement ranging from $18,000 to $400,000 for what sounds like the same work. The spread is real, and it is almost never about the model. Here are the ranges, the assumptions under them, and the structure that keeps the money honest.

What a proof of value is actually for

A proof of value answers two questions and no others: does this work on your data, and is the production version worth buying? Everything that does not serve those two questions belongs in a later phase. That sounds obvious, and it is the single most common place a first engagement goes wrong. A scope that also promises a user interface, an authority to operate, three integrations and a training program is not a proof of value. It is a small production build with a pilot label on it, priced like a pilot and delivered like a pilot, which is to say late.

The distinction matters because the two documents behave differently under pressure. A production build gets judged on whether it shipped. A proof of value gets judged on whether the answer it produced was trustworthy. When a program office or a commercial buyer asks us to size a first engagement, we separate those in writing before anyone talks about price.

The second thing we do is write down the decision the engagement is supposed to inform. "We want to see what AI can do here" is not a decision. "We will fund a production system in FY27 if a model can extract these fourteen fields from these forms with a false-extraction rate under four percent, measured on records nobody trained on" is a decision. The second version costs less to run than the first, because scope stops expanding the moment the acceptance metric is fixed.

The scope that fits inside a first engagement

A well-formed proof of value has one workflow, one data source, one metric and one environment. Each additional item on any of those four axes multiplies the work rather than adding to it, because every combination is evaluated separately. Two data sources with two environments is four integration paths.

What fits: a single decision or document type, real records from the customer, a measured baseline, one iteration on the errors, and a held-out evaluation that produces a number. What does not fit: production hardening, Section 508 conformance (36 CFR Part 1194), an ATO package, role-based access control, a full user interface, and integration into the system of record. Those belong in the follow-on, and estimating their cost is itself a deliverable of the proof.

  • One workflow: a single decision, document class or prediction, described in the customer's own vocabulary.
  • Real records: the customer's data, including the ugly ones, not a vendor sample set.
  • A pre-registered metric: the number that decides the outcome, written down before any results exist.
  • A non-AI baseline: current throughput, current error rate, or a rules-based comparison the model has to beat.
  • One environment: the place the work runs, named specifically, with its access constraints known on day one.
  • A production estimate: cost, timeline and blockers for the real build, produced from what the proof learned.

What it costs, and the assumptions under each number

Prices below assume a small team of senior people working part-time across the engagement rather than a large team working full-time. That structure is deliberate. A first engagement is mostly thinking, error analysis and conversation with subject-matter staff, and those activities do not compress by adding bodies. A useful rule for sanity-checking any quote: a fully burdened senior engineering rate in this market runs roughly $180 to $260 per hour, so a six-week engagement with two people at half time is on the order of 240 hours and lands between $50,000 and $70,000 of labor before infrastructure and travel.

Scope tierDurationTypical priceWhat it can answer
Data readiness sprint2–3 weeks$18K–$35KWhether the data can support any model at all, and what it would take to fix it. Run this first when nobody can describe the data's completeness.
Narrow feasibility check3–4 weeks$25K–$45KOne question, one data source, one metric. Good when the data is already clean and the decision is binary.
Standard proof of value6–8 weeks$60K–$120KEnd-to-end path on real records, a measured baseline, one iteration on the errors, and a costed production plan.
Regulated or disconnected10–12 weeks$120K–$250KThe same result inside CUI, HIPAA, CJIS or air-gapped constraints, where access approval alone consumes weeks.

Two things move a quote outside these bands legitimately: a hardware or sensor component, which brings procurement lead times that software estimates do not cover, and a security environment where the vendor's engineers cannot reach the data and every experiment runs through cleared staff on the customer's side. Anything else that pushes a first engagement past $250,000 deserves a line-item explanation, and a buyer is entitled to ask for one.

Data readiness drives the cost, not the model

Model selection is a decision our engineers make in an afternoon. Getting a legally clean, representative, labeled extract of production data into an environment where work can happen is what takes six weeks and eats the budget. That inversion surprises buyers who assume they are purchasing algorithm work.

What drives the price of a first engagement

Data access and data readiness
92%
Clarity of the acceptance metric
86%
Where the work has to run (cloud, on-prem, air-gapped)
79%
Subject-matter time available for labeling and review
76%
Number of stakeholders who must agree on "good"
70%
The model and infrastructure work itself
63%

Relative weighting from practitioner reading of first-engagement scoping. Illustrative, not a measured statistic.

Three data facts change a price more than any technical choice. Whether an extract already exists that someone is authorized to hand over. Whether ground truth exists, or whether the engagement has to create it from expert review. And whether the records carry a legal control regime: CUI under 32 CFR Part 2002 and NIST SP 800-171, protected health information requiring a business associate agreement under 45 CFR 164.504(e), criminal justice information under the FBI CJIS Security Policy, or student records under FERPA at 34 CFR Part 99. Each regime adds review time before a single record moves.

A buyer who wants a cheaper proof of value has one reliable lever, and it is not negotiating the rate. It is arriving with a documented extract, a named data owner with authority to approve release, and eight to twelve hours per week of a subject-matter expert's time. That combination routinely cuts a quote by a quarter, because it removes the largest source of schedule risk the vendor is pricing against.

A proof of value that can only end in "yes" is a demo with an invoice attached.

Why six to eight weeks is the honest answer

Shorter engagements are worth buying, but a three-week check answers only a narrow question on clean data. Eight weeks is the workhorse duration because the sequence below has a hard critical path: nothing after step two starts until real records are in hand, and error analysis cannot begin until there are errors to analyze.

Standard proof-of-value sequence

1
Acceptance metric and threshold signed. Data owner named. Environment chosen.
Week 0–1
2
Real records landed. Ground truth built with subject-matter review. Non-AI baseline measured.
Week 1–2
3
First end-to-end run against production-representative records. Number on the board, however bad.
Week 2–4
4
Error analysis on the failure cases. One focused iteration, not five speculative ones.
Week 4–6
5
Held-out evaluation run once, on data untouched all engagement. Result frozen.
Week 6–7
6
Readout, code and evaluation suite handover, production cost estimate with assumptions.
Week 7–8

Step two is the one most often skipped, and skipping it is fatal. A team that never measures the current process cannot say whether the model helped. If clerks process 40 forms an hour at a two percent error rate, that is the number to beat, and an afternoon of stopwatch work establishes it. Without that baseline, an impressive accuracy figure can still be worse than what the customer already has.

Structuring payment against milestones

Firm-fixed-price under FAR 16.202 is the right instrument for a proof of value, because the scope is bounded and the customer should not carry the vendor's estimating risk. Time-and-materials on a first engagement transfers every schedule surprise to the buyer, and schedule surprises are guaranteed in the data-access phase.

Fixed price does not mean payment on completion. Split the price across four events tied to observable artifacts, using performance-based payments under FAR Subpart 32.10 on a federal contract, or plain milestone invoicing on a commercial agreement:

Milestone 1, 20% at data landing. Paid when the agreed extract is in the working environment and the acceptance metric is signed by both sides. It sits early because the vendor cannot control how long the customer's approval chain takes.

Milestone 2, 25% at baseline. Paid when ground truth exists and the non-AI baseline is measured and documented. At this point the customer already owns something useful: a labeled evaluation set that outlives the engagement.

Milestone 3, 30% at first end-to-end result. Paid when the system runs the whole path on real records and produces a measured number against the metric. The number does not have to be good. It has to be real and reproducible.

Milestone 4, 25% at handover. Paid on delivery of the frozen held-out evaluation, the source code, the evaluation suite, the written findings and the production estimate. Withholding a quarter of the price against the final package is the mechanism that guarantees the documentation actually gets written.

Two clauses belong alongside the payment schedule. A stop-work trigger: if the data has not landed by an agreed date, either party can halt and settle at milestones earned, so a stalled engagement does not burn its budget waiting. And a rights position stated up front, whether that is government purpose rights under DFARS 252.227-7014, the data clause at FAR 52.227-14 for a civilian agency, or a commercial license grant. Rights arguments discovered at closeout are expensive and avoidable.

What ships at the end regardless of the answer

The engagement can conclude that the approach does not work. That is a legitimate and valuable outcome, and it does not reduce the delivery obligation. Everything below should be in the customer's hands at closeout whether the verdict was go or no-go.

  • Source code in a repository the customer controls, with the rights position written into the agreement, not asserted afterward.
  • The cleaned, labeled evaluation set plus the code that assembled it, which is often worth more than the model.
  • A re-runnable evaluation suite so the next vendor, or the customer's own staff, can reproduce the number without trusting anyone.
  • The measured baseline, including current-process throughput and error rate.
  • Written findings including the negative results: what failed, on which record types, and the hypothesis for why.
  • A production cost and schedule estimate with its assumptions exposed, so it can be defended in a budget review.
  • The blocker list: data gaps, access obstacles, policy questions and integration surprises found along the way.

That last item often justifies the whole engagement. Finding three unmapped data-quality problems and one unresolved records-retention question saves a production program from discovering them at a far worse moment. Our teams write the blocker list as a standing deliverable for that reason.

A proof of value that cannot fail is worthless

If the engagement is structured so that every possible result is reported as success, it has no information content and the money bought nothing. A proof of value that can only end in "yes" is a demo with an invoice attached. The failure modes are recognizable, and a buyer can screen for all of them before signing.

The metric gets chosen after the results are in. The evaluation runs on the same records used to build the system. The demonstration is a curated set of favorable examples. There is no threshold, only a narrative. Or the vendor holds the code, so the number cannot be independently checked. Each converts a measurement into a sales artifact.

The pre-registered decision rule

Write the sentence that would end the project

Before any modeling starts, both sides sign one sentence of this shape: "If the false-extraction rate on the held-out set exceeds four percent, or median review time per record does not fall below ninety seconds, we do not recommend proceeding to production." The threshold is set from the operational requirement, not from what looks achievable. Then the held-out set is run once, at the end. A buyer who insists on this one paragraph has removed most of the ways a first engagement can waste money.

Where the money comes from on the federal side

Federal buyers have several routes to fund a first engagement, and the price bands above map onto them cleanly. Simplified acquisition procedures under FAR Part 13 cover most proof-of-value work, since an engagement priced under the simplified acquisition threshold in FAR 2.101 avoids the full FAR Part 15 source-selection burden. Purchases of commercial services proceed under FAR Part 12 with the streamlined terms at FAR 52.212-4.

For agencies that already hold a schedule, IT professional services under GSA Multiple Award Schedule SIN 54151S is the shortest path from decision to kickoff. Program offices with prototype authority can use an other transaction under 10 U.S.C. 4022, or a Commercial Solutions Opening under 10 U.S.C. 3458, both of which are designed for exactly this kind of bounded experiment. And an SBIR Phase I is, in substance, a government-funded proof of value: awards commonly run from $150,000 to roughly $300,000 depending on the agency, with the resulting SBIR data rights protection period now set at twenty years under the SBIR Policy Directive.

State and local buyers have their own version of the same logic. Most states allow an informal quote process below a statutory small-purchase threshold, and cooperative purchasing agreements let one agency ride another's competitively awarded contract. Ask the contracting office which route makes a bounded eight-week engagement easiest before writing a statement of work that forces a full solicitation.

Reading a quote you have already received

Most buyers arrive holding two or three proposals that are hard to compare. These checks separate them fastest.

The price has no line for data work

If the largest driver of cost does not appear in the cost breakdown, the vendor either has not thought about it or intends to raise it as a change order later. Ask how many hours are budgeted for data access, extraction and ground-truth construction, and what happens to the price if the extract arrives three weeks late.

The final deliverable is a presentation

A slide deck is a summary of deliverables, not a deliverable. If the closeout package does not include running code, the evaluation set and code the customer can execute without the vendor, the engagement leaves nothing behind that survives the relationship.

The metric appears for the first time in the final report

A metric selected after results exist is a metric selected to flatter the results. The acceptance criterion and its threshold belong in the statement of work, signed before the first model run.

The demonstration runs on the vendor's sample data

Performance on a public benchmark predicts very little about performance on your records, which carry their own scanning artifacts, legacy formats, local abbreviations and decades of accumulated exceptions. Insist the reported number comes from your data.

The rights position is silent

Silence in the agreement is resolved in the vendor's favor at closeout. Name the clause in the document: government purpose rights, unlimited rights, or a specific commercial license with its restrictions written out.

Bottom line

A first AI engagement should be small enough to fund without a fight, long enough to touch real data, and structured so a negative result is reported as clearly as a positive one. Six to eight weeks and $60,000 to $120,000 fits most software-shaped problems on reasonably available data, with a cheaper tier for clean narrow questions and a longer one for regulated or disconnected environments. Pay against four observable milestones. Fix the metric before the work starts. Take delivery of the code and the evaluation set no matter which way the answer lands. That is a real answer for a known price, which is the entire point of the exercise.

Frequently asked questions

How much should a first AI engagement cost?

Most software-shaped proofs of value land between $60,000 and $120,000 over six to eight weeks. A narrow feasibility check on clean data runs $25,000 to $45,000 in three to four weeks. Regulated or disconnected environments run $120,000 to $250,000, because access approval consumes weeks before modeling starts.

What makes one AI pilot quote three times another?

Almost always the data assumptions, not the technical approach. A vendor assuming a clean extract arrives on day three prices very differently from one assuming it must build ground truth from expert review inside a controlled environment. Compare the data-work hours before comparing the totals.

Should a proof of value be fixed-price or time-and-materials?

Firm-fixed-price under FAR 16.202, split across four milestones. The scope is bounded, so the vendor should carry the estimating risk. Add a stop-work trigger if the data has not landed by an agreed date, settling at milestones earned.

What should we receive if the pilot concludes the approach does not work?

Everything except a production recommendation: source code, the cleaned and labeled evaluation set, a re-runnable evaluation suite, the measured baseline, written findings covering the negative results, a production cost estimate, and the blocker list. The evaluation set alone often outlives the engagement by years.

How do we make a first engagement cheaper without cutting scope?

Arrive with an approved data extract, a named data owner with release authority, and eight to twelve hours per week of subject-matter expert time. That combination removes the largest schedule risk in the vendor's estimate and reliably reduces the price, which negotiating the hourly rate does not.

1 business day response

Sizing a first AI engagement?

Send us the workflow and what you know about the data. Our team will come back with a scoped range, the acceptance metric we would propose, and the milestone structure behind it.

Start a conversationHow we workMore insights →
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE