Skip to main content
Commercial

A 30-day AI pilot that proves something

Most pilots end with a demo, a deck of favorable examples, and an argument about whether it worked. A pilot that ends in a decision is built differently: the claim is written before kickoff, the baseline is measured before anything is built, the test data stays on the buyer's side, and the result is allowed to come back no.

Three pilots, no decision

An executive who has run three AI pilots and bought nothing usually did not hire three bad vendors. Each engagement ended the same way. Something ran. A deck showed twelve examples where the output looked good. Somebody asked whether this was better than what the company already does, and the room discovered that nobody had written down what the company already does, or how good the new thing had to be, or who would decide. The engagements were not experiments. They were exhibitions, and an exhibition has exactly one possible outcome, because the party building it also chooses the examples, the metric, and the moment to stop.

Thirty days is enough time to run a real experiment on a narrow question. It is not enough time to survive a vague one. The difference is almost entirely design, and almost all of the design happens before day one. What follows is the structure that turns a month of work into a result a finance committee can act on, whether the answer is yes or no.

What makes a day-30 result decision-grade

A named decision that changes based on the result
95%
Current process measured on the same items
92%
Test items chosen and held by the buyer
89%
Enough scored items to separate signal from noise
83%
Scoring rule fixed in writing before scoring starts
78%
Failure behavior characterized, not just accuracy
71%

Editorial weighting from published federal guidance and delivery practice. Illustrative, not a measured statistic.

A demo is an existence proof. A pilot is a measurement.

A demo answers the question can this ever work. For most language and document tasks in 2026, the answer is yes, and the cost of obtaining that answer is close to zero. It tells a buyer nothing about the purchase, because nobody is deciding whether the technology can ever produce a correct output. They are deciding whether it produces correct outputs often enough, on their own messy inputs, to change a process that already runs.

A measurement answers four questions at once. How often is it right on work that looks like our work. How does that compare to what we do now, measured the same way. What does it do when it is wrong, and can a person catch it. What does it cost when it runs every day rather than once in a conference room. Those questions have numbers attached, and numbers can be argued with, which is the point.

The two look identical in the room. Same laptop, same screen, same competent engineer explaining the output. One structural test separates them: who picked the examples. If the party building the system also selected what it was shown, the meeting revealed the builder's judgment about which cases flatter the system. That is real information about the builder, and none about the system.

Write the decision first, then the sentence the pilot has to earn

Before scoping any technical work, write four things on one page: the decision, the person who owns it, the date it gets made, and the two branches. "By October 6, the VP of Operations decides whether to fund a production build for the claims intake queue, or to stop and keep the current process." A pilot without that paragraph is a research budget wearing a business costume.

Then write the claim as one sentence with no adjectives in it. The shape: on this task, the system reaches this metric at a level of at least this threshold, measured on this many items drawn from this population, scored by this person under this written rule. Every blank has to be fillable at kickoff. If the threshold cannot be filled in, the organization has not decided what would be good enough, and no result will satisfy it. If the population cannot be named, the pilot runs on whatever data is easiest to extract, which is never the data the decision is about.

The threshold comes from economics, not ambition. Work out what the current process costs per item, what the new one would cost including review, and what an error costs. The threshold is the accuracy at which the change pays for itself with room for the errors nobody predicted. A number picked because it sounds impressive produces a pilot that lands at ninety-one percent and starts an argument instead of ending one.

Write the losing sentence too. "If the system scores below this, we stop, and here is what we do with what was built." Teams resist that line because it looks like planning to fail. It is the opposite: it is the only thing that makes the yes mean anything.

The baseline is measured, not remembered

Ask any operations leader what the current error rate is on a manual process and the answer will be a number stated with confidence and derived from nothing. Remembered baselines run optimistic, because the memorable cases are the ones somebody caught. The fix is cheap: take the same sample of items the pilot will use, run them through the existing process, and score them under the same rule. On most document and triage workflows this is half a day to two days of effort.

Two outcomes happen often enough to plan for. The measured baseline lands far worse than anyone believed, and a modest system clears a bar the organization thought was high. Or it lands near-perfect on the cases that matter, which means automation should be aimed at throughput and cost rather than quality, and the pilot's metric was wrong. Both are worth two days, and both are invisible without them.

Federal buyers are being pushed toward the same discipline. OMB Memorandum M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government, issued April 3, 2025, tells agencies to consider whether performance metrics are correctly tied to desired outcomes and whether the agency "can adequately measure baseline performance" before writing incentives against them. Its companion, M-25-21, requires that an AI impact assessment document the intended purpose and expected benefit "supported by specific metrics or qualitative analysis," assessed "as compared to existing agency processes." Both memoranda were signed the same day and replaced the prior guidance, M-24-18 and M-24-10 respectively.

The holdout belongs to the buyer

This is the single design choice that separates the two kinds of pilot, and it is the one most often given away without a fight. The test set has to be selected by the buyer, held by the buyer, and never visible to the party building the system until the scored run.

The federal acquisition guidance is unusually direct here, and the language transfers cleanly to a commercial contract. M-25-22 instructs agencies to use data they have defined, such as agency validation and testing datasets, when conducting independent evaluations, and states that the data used "should not be accessible to the vendor, and should be as similar as possible to the data used when the system is deployed." It goes on to require that vendors provide the access and time necessary for the agency to complete an independent evaluation; that where a vendor performs the testing, results be detailed enough for the testing to be independently verified or reproduced; and that contracts detail the examination, testing, and validation procedures and not prohibit the agency from internally disclosing how the vendor conducts testing or the results of that testing.

Three clauses, and they are the whole game. The buyer defines the evaluation data. The vendor cannot see it. The result has to be reproducible by someone else. A commercial pilot agreement carrying those three sentences produces a different engagement than one without them, whoever the vendor is.

The mechanics take an afternoon. Draw the sample before kickoff, stratified so it includes the ugly cases and not only the clean ones. Freeze it and keep the answer key. Hand over inputs only. Draw the development slice separately, and check that no item, no near-duplicate, and no downstream artifact of a holdout item appears in it. Then leave it alone until week four.

A pilot that cannot come back with a no is not a pilot. It is a purchase with a waiting period.

How many items it takes to tell a difference from noise

The most common quiet failure of a 30-day pilot is a real-looking number computed on forty examples. Forty examples cannot distinguish an eighty-five percent system from a ninety-two percent system. The result is not wrong so much as it is empty, and it will be defended for months by whoever likes it.

The table below gives the approximate number of scored items needed per arm to detect a stated improvement over a baseline, using the standard normal-approximation sample size for comparing two proportions, at a five percent two-sided significance level and eighty percent power, without a continuity correction. "Per arm" means the baseline process and the new system are each scored on that many items.

Improvement to be shownScored items per armTotal items scoredFeasible in 30 days
70% → 90% (20 points)About 62About 124Yes, comfortably
80% → 90% (10 points)About 199About 398Yes, if scoring is cheap
85% → 90% (5 points)About 686About 1,372Only with automated or pre-labeled scoring
90% → 93% (3 points)About 1,356About 2,712Rarely, at expert-review cost

Read that as a scoping constraint rather than a statistics lesson. A 30-day pilot can prove a large improvement and cannot prove a small one. If the business case rests on three points of accuracy against an already-strong process, either the budget covers thousands of expert-scored items or the pilot has to be aimed at a different claim: cycle time, cost per item, coverage of a queue nobody currently touches, or the error types that cause the expensive failures. Those are often the honest claims anyway.

One design choice buys back much of that cost. Run both the current process and the new system over the same items, then compare them item by item and count only the cases where they disagree. Per-item difficulty cancels out, the comparison becomes far more efficient than two independent samples, and the analysis reduces to which side wins the disagreements. When a pilot has to be small, paired scoring on identical items is usually the only way to make it mean anything.

Report the confidence interval, not the point estimate. "Ninety-one percent" is a story. "Ninety-one percent, with a 95 percent interval from 86 to 95, against a measured baseline of 78" is a decision. If the interval overlaps the threshold, the honest report says the pilot did not resolve the question and states how many more items would.

The thirty days

The shape below assumes one narrow task, one integration point, and data access that has already been approved. Data access is the dependency that most often eats a pilot alive, which is why it sits before day one rather than inside week one.

Thirty days, counted from an approved data path

1
Decision, claim sentence, threshold, scoring rule; holdout drawn and frozen
Before kickoff
2
Access executed; current process scored on the holdout to set the baseline
Days 1 to 4
3
Build and iterate on the development slice only
Days 5 to 14
4
Dry run on a small internal set, error review, version frozen and recorded
Days 15 to 21
5
Single scored pass on the holdout; no tuning, no second attempt
Days 22 to 25
6
Analysis, cost model, written recommendation, decision meeting on the calendar
Days 26 to 30

Week three is the week teams want to skip and the week that saves the pilot. A dry run on internal data surfaces the boring failures: a date format that breaks parsing, a document class nobody mentioned, an API rate limit, a scoring rule two reviewers interpret differently. Every one of those is cheap on day 18 and fatal on day 24.

One pass on the holdout, and the version is frozen

The most common way an honest pilot goes wrong is tuning against the holdout. Someone runs the scored set, sees eighty-four percent, spots a fixable pattern in the errors, fixes it, and runs again. Nobody is being dishonest, and the number is inflated anyway: the test set has been partly absorbed into development, and the second result reflects fit to those particular items rather than performance on new ones.

Hold the line with mechanics rather than willpower. Fix the split at the start and do all iteration on the development slice. Before the scored run, freeze and record the model identifier and version, the prompt or configuration, the retrieval index and its build date, any thresholds, and the preprocessing code, with a hash written into the pilot record. Run once. If the result disappoints, the honest move is a second, separately drawn holdout for a later round, not a rerun on the same items.

Version discipline is not a research nicety. The same federal acquisition guidance tells agencies to require vendors to meet performance standards before deploying a new version of an AI system or service, and to roll back to a previous version if a new one fails. If measured behavior is allowed to move underneath a contract, the measurement it was bought on no longer describes what is running.

Three ways a careful pilot still misleads

Leakage. Something in the input silently contains the answer. A resolved ticket carries the resolution code in a metadata field. A scanned form includes a reviewer's stamp. A "raw" export was quietly filtered to the records that already had clean labels. Leakage produces suspiciously excellent results, and the cheapest detector is suspicion itself: when a first run comes back near-perfect, stop and go looking for the leak before celebrating.

The easy slice. The pilot ran on the vendor-facing document type with consistent formatting, while the real queue is forty percent handwriting, faxes, and a legacy system that emits fixed-width text. A result on the easy slice generalizes to the easy slice. Stratify the holdout to match the production mix, and report accuracy per slice rather than one blended number, so the decision can be scoped to the slices where the system actually works.

Scoring by the party that built it. Judgment calls drift toward the scorer's expectation, and the effect does not require bad faith. Have someone independent score the run against a written rule, blind to which output came from which system wherever outputs can be anonymized, and measure agreement between two scorers on a subset. If two reasonable people disagree on a fifth of the items, the metric needs fixing before it can settle anything. M-25-21 makes the same instinct a federal requirement: an independent reviewer within the agency who was not involved in development identifies concerns or gaps, and a named individual signs the risk acceptance.

What ships on day 30 whatever the answer

A pilot that produces a no should still leave the organization better off than it was on day one. That is what distinguishes an experiment from a sunk cost, and it is a reasonable thing to require in the agreement.

  • The frozen holdout with its answer key, reusable against the next vendor, the next model version, or the next internal attempt.
  • The measured baseline, often the most durable output, because it prices the current process for the first time.
  • Scored results with confidence intervals, broken out by data slice and error type rather than one blended number.
  • An error taxonomy naming what the system got wrong and how a person would catch each class in production.
  • The running cost model: cost per item at pilot volume and at full volume, including human review and monitoring.
  • The code, configuration, and data pipeline, with rights that let the buyer or a different firm pick it up.
  • A written recommendation with the reasoning attached, including what a second round would have to test.
Design elementPilot that proves nothingPilot that proves something
Test dataExamples selected by the party building the systemStratified sample drawn and held by the buyer, unseen until the scored run
BaselineAn estimate someone recalls in a meetingThe current process scored on the same items under the same rule
Sample sizeHowever many fit in the deckSized to the improvement the business case requires
MetricAgreed after the results are inWritten with the threshold before kickoff, tied to cost per item
IterationRepeated runs until the number looks goodOne scored pass on a frozen version, hashes recorded
OutcomeA recommendation to keep exploringA dated decision, with a no that is allowed to happen

The federal version of the same rules

Buyers inside government are working from published policy that reads like a specification for the pilot described above, which makes it useful to commercial readers too. OMB M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, dated April 3, 2025, defines high-impact AI as AI whose output serves as a principal basis for decisions or actions with a legal, material, binding, or significant effect on rights or safety, and attaches a set of minimum risk management practices to those uses. The first practice on the list is pre-deployment testing that reflects expected real-world outcomes.

The memorandum also carves out pilots, and the conditions on the carve-out are instructive. A pilot program for a proposed AI use case is exempt from the minimum practices provided the program is of limited scale and duration, the agency Chief AI Officer has certified that it may go forward with the certification tracked centrally, individuals who may interact with the AI can opt in and out where possible with sufficient notice, and the minimum practices are applied where practicable. Limited scale and limited duration are the price of the exemption. A tightly scoped 30-day engagement fits that shape; a nine-month "pilot" quietly running in production does not.

On the acquisition side, M-25-22 directs agencies to seek demonstrations and tests "in scenarios that closely reflect the intended real-world operating environment," to use them to interrogate both the capabilities and the limitations of a provider, and to test proposed solutions during evaluation to the greatest extent practicable. It also pushes performance-based techniques, quality assurance surveillance plans, protections against vendor lock-in through knowledge transfer and data and model portability, and sunset criteria for when an agency should reconsider continued use.

For a consultancy, integrator, or prime bringing a technical partner to a client, the practical read is short. A partner who arrives with a written claim, a buyer-held holdout, an independent scoring rule, and a report that survives being reproduced reduces risk on the part of the engagement the client scrutinizes hardest. A partner who arrives with a demo transfers that risk to whoever sponsored them.

One mechanism raises the stakes on getting the criteria right. Under 10 U.S.C. 4022, a Defense Department prototype other transaction can be followed by a production contract or transaction awarded without further competition, provided competitive procedures were used to select the participants and the participant successfully completed the prototype project. Successful completion is judged against the success criteria written into the agreement, which makes that paragraph the most valuable text in the document, and it is written before any work begins.

Where the standards sit is worth stating plainly, because they are moving. The NIST AI Risk Management Framework 1.0, released January 26, 2023, organizes work around four functions, Govern, Map, Measure, and Manage, and its Generative AI Profile, NIST-AI-600-1, followed on July 26, 2024. NIST has said the framework is being revised as part of the White House AI Action Plan, so a contract written today should cite the version and date rather than the framework in the abstract.

What the pilot agreement should say

Fixed price, fixed end date, fixed scope, and no automatic conversion. A pilot on time and materials with an open end has no forcing function, and the decision slips until the sponsor loses interest. A pilot that rolls into a production subscription unless cancelled is a sales motion rather than an experiment. Keep the two decisions separate on paper.

Custody of the evaluation data. Name who draws the sample, who holds it, when it is released, and that the party building the system does not receive it before the scored run.

Who scores, and against what rule. The scoring rule is an attachment, not a conversation. Include the tie-break procedure and the agreement check between scorers.

Version freeze and disclosure. The configuration under test is recorded and cannot change mid-evaluation, and any change afterward triggers a re-measurement rather than an assumption of continuity.

Deliverables that survive a no. List the seven items above and make them due on the end date regardless of outcome. This is the clause that turns a failed pilot into a priced, documented decision.

Data return, deletion, and rights to what was built. Say what happens to the buyer's data at the end, in what format it comes back, on what timeline it is deleted, and whether it may be used for any model training. The buyer should also be able to hand the pilot artifacts to a different firm without renegotiating, which federal guidance frames as vendor lock-in protection through knowledge transfer, data and model portability, and clear licensing terms.

Bottom line

A pilot is an instrument. Its job is to produce a number that a reasonable skeptic, holding the same data and the same scoring rule, would produce again. Everything that makes a pilot feel productive in week two, the polished interface, the impressive examples, the enthusiasm in the standup, is irrelevant to that job and occasionally hostile to it. Write the claim, measure the baseline, keep the test set, size the sample to the difference that matters, freeze the version, score once, and put the decision on the calendar before anyone starts building. That is thirty days well spent, and it is the same thirty days whether the answer turns out to be yes or no.

Frequently asked questions

Is 30 days long enough to evaluate an AI system?

For one narrow task with an approved data path, yes. Thirty days is enough to measure a baseline, build on a development slice, and score a single frozen run on a held-out sample. It is not enough to evaluate a platform, integrate three systems, or show a small improvement over an already-strong process. The binding constraint is not calendar time. It is how many items can be scored and how large a difference has to be shown.

Why should the buyer hold the test data instead of the vendor?

Because whoever holds the test data controls the result. Federal acquisition guidance in OMB M-25-22 states that data used for independent evaluation should not be accessible to the vendor and should be as similar as possible to the data used when the system is deployed, and requires that results be detailed enough to be independently verified or reproduced. The same arrangement works commercially: the buyer draws and holds a stratified sample, hands over inputs only, and keeps the answer key until the scored run.

How many test items does a pilot need?

It depends entirely on how large an improvement has to be shown. Using the standard sample size calculation for comparing two proportions at five percent significance and eighty percent power, showing a jump from 70 to 90 percent takes roughly 62 scored items per arm, 80 to 90 percent takes roughly 199, and 85 to 90 percent takes roughly 686. Scoring both systems on the same items and comparing them pair by pair is substantially more efficient than two independent samples, and is usually the right design for a short pilot.

What should a pilot deliver if the answer is no?

The frozen evaluation set with its answer key, the measured baseline, scored results with confidence intervals broken out by slice and error type, an error taxonomy, a running cost model, the code and pipeline with usable rights, and a written recommendation. Those artifacts keep their value for the next attempt, and requiring them on the end date regardless of outcome is what makes a no cheap enough to accept.

Do federal buyers treat pilots differently from production systems?

Yes. OMB M-25-21 exempts a pilot program for a proposed AI use case from the minimum risk management practices required of high-impact AI, but only if the program is of limited scale and duration, the agency Chief AI Officer certifies it with the certification tracked centrally, affected individuals can opt in and out where possible, and the minimum practices are applied where practicable. The exemption is built for short, bounded evaluations, not for long-running deployments labeled as pilots.

1 business day response

Need a pilot that ends in a decision?

We design and run 30-day evaluations on real data: written claim and threshold, measured baseline, a holdout you keep, a single scored pass on a frozen version, and a report with the confidence intervals attached.

How we workMore insights →Start a conversation
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE