The most expensive outcome of a federal pilot is not failure. It is a pilot that works, delights the people who ran it, produces a written result everyone agrees is good, and then goes nowhere for eleven months because nobody decided in advance how the government would buy the thing. That pattern is so common it should be treated as the default outcome of an unplanned pilot. The pilot is not a test of the product. It is the first phase of a purchase, and if it was not designed as one, the purchase does not follow.
This is written for the person carrying the public sector number at a commercial software company: how to design a pilot so that its success has somewhere to go, what has to be decided before it starts, what has to be produced during it, and what the expansion story has to look like when it ends.
Why good pilots die
Five causes, and only one of them is technical.
No purchase path was chosen before the pilot began. The pilot was funded from a discretionary pot or run under an existing contract with room in it. When it succeeds, the follow-on needs a different instrument, and choosing one takes months of work that nobody started. Meanwhile the fiscal year closes.
The sponsor was the wrong person. The pilot was run by an innovation office, a research group or a technical staff whose mandate is to try things. Trying things is their job; buying things is not. The result lands in a report that a program office may or may not read.
The metric was interesting rather than budgetary. The pilot proved the product is accurate, or fast, or well liked. None of those is a line in anybody's budget justification. The number that moves a purchase is one the program already reports to someone above it.
The security work was deferred. The pilot ran under a temporary authorization, in a sandbox, on synthetic or de-identified data. Production requires a full authorization that nobody started, so success adds nine months instead of subtracting them.
No expansion story existed. The pilot covered one office and one workflow. Nobody wrote down what the same system does across the whole program, what it costs at that scale, and what has to be true for it to get there. So the agency has a proven point solution and no basis for a larger commitment.
What determines whether a pilot converts to a purchase
Editorial weighting, illustrative rather than measured. The last row is low because breadth of demonstration rarely moves a purchase and often complicates the scope.
Decide five things before the pilot starts
These are cheap to settle in week zero and nearly impossible to settle in the final week, and each one is a conversation with a different person.
One: the sponsor who can buy
Identify the person whose budget would carry the enterprise purchase and whose mission outcome the system improves. That person may not be the one who invited you in. If the pilot's enthusiast sits in an office with no acquisition budget, the pilot needs a second sponsor from the start, and getting one is part of the pilot's design rather than a step after it. A pilot whose only sponsor cannot buy is a demonstration with a longer timeline.
Two: the instrument
Ask the contracting officer directly, before the pilot begins, what instrument the follow-on would ride and what its ceiling is. This is a normal question and contracting officers generally answer it. The answer determines a great deal: whether the purchase can be made without competing it, whether an existing vehicle covers it, whether your company or a reseller needs to hold something it does not currently hold, and how long the process takes. If the answer is that no path currently exists, that is the most valuable thing you will learn all quarter, and there is time to build one.
Three: the measurement
Write the success criteria as a number the program already tracks, with a baseline measured before the pilot starts. Cycle time on a process the program reports. Backlog size. Error rate on a review that has a stated target. Cost per transaction. The baseline matters as much as the result, and it is only obtainable before you start. A pilot that ends with a good result and no baseline produces an argument rather than evidence.
Four: the data and the environment
Decide whether the pilot runs on real data in the environment the production system will use, or on a copy somewhere convenient. The convenient version is faster to start and much slower to convert, because it proves the model and proves nothing about deployability, and deployability is what the enterprise decision turns on. Where a full production environment is not available for a pilot, run in the same class of environment with a written path to production stated at the outset.
Five: the enterprise scope
Before the pilot begins, write the thing the pilot is a step toward: what the system does at full program scale, how many users, what integrations, what it costs, what has to be true for it to run there. Not a proposal, a page. Its purpose is to make the pilot a first increment of something rather than a self-contained experiment, and to give the sponsor language for a conversation with a budget owner while the pilot is still running.
What the pilot has to produce besides a result
A pilot designed to convert produces four artifacts, and three of them have nothing to do with whether the product worked.
| Artifact | Produced by | What it does for the purchase |
|---|---|---|
| Measured result against a baseline | Instrumentation inside the running system | Gives the sponsor a defensible number for a budget justification |
| Security documentation package | Engineers, during the pilot, not after | Removes the largest post-pilot delay from the critical path |
| Deployment definition and runbook | The pilot deployment itself, if built properly | Makes the enterprise rollout an execution task rather than a new project |
| Written enterprise scope and price | The vendor, reviewed with the sponsor mid-pilot | Lets the sponsor start the funding conversation before the pilot ends |
The second row is where most of the recoverable time is. Security authorization is the longest pole in almost every federal software purchase, and it is the one most often started after a successful pilot rather than during one. Doing the documentation work concurrently with the pilot, using the pilot deployment as the evidence, can remove two or three quarters from the path to production. It requires engineers who can write the artifacts accurately while they build, which is a staffing decision made at the start.
Designing the pilot as a first increment
The engineering shape of a convertible pilot differs from a demonstration in specific ways.
Build it where production will live. The same cloud region, the same identity provider, the same network posture, the same logging destination. If the production destination is not yet available, build against its documented constraints so that migration is configuration rather than rework.
Instrument the outcome, not the usage. Product analytics tell you people clicked. The pilot needs the program's own metric computed continuously from the system's data, in a form the program office can reproduce. Build that measurement as a feature of the pilot deployment and give the sponsor a page they can look at.
Log decisions from the first day. For each action the system influences, record the version, the inputs or a hash where inputs are sensitive, the output, and whether a person overrode it. This is what supports the result at the end and what a security reviewer will want anyway. Retrofitting it is nearly impossible because the information exists only at the moment of the decision.
Make the deployment reproducible. Infrastructure as code, a single deployment manifest, and an install a government engineer can run without vendor credentials. If the enterprise rollout means repeating a manual setup in nine offices, the rollout gets estimated pessimistically and the purchase shrinks.
Handle accessibility before users touch it. Government users include people who use screen readers and keyboard-only navigation, and the requirement is part of what the government may buy. A pilot interface that fails an accessibility review generates a finding that follows the product into the purchase.
Scope narrowly and finish. One workflow, done completely, beats three workflows demonstrated partially. The enterprise argument is made by depth on a real process, because that is what a program office can extrapolate from.
The four ways a pilot ends, and what each one is worth
Sales forecasts treat a pilot as a binary. In practice there are four endings, and three of them are recoverable if you can tell which one you are in while it is happening.
It converts. The result met the criteria, the instrument existed, the security work was done or nearly done, and the sponsor had a priced scope to take to a budget owner. The follow-on starts within a quarter or two. This is the outcome the design produces, and its frequency is a function of the five decisions at the start rather than of how good the technology was.
It converts smaller and later. The result was good and something structural was missing: the instrument had to be built, or the budget line sat in a different program, or the authorization work was not started. The purchase happens, at a reduced first increment, one or two fiscal years out. This ending is common and it is not a loss, but it costs the company the carrying period and often costs the sponsor their enthusiasm. Every one of the delays is preventable at the start.
It becomes a repeated pilot. The result was good, nobody could buy, and the following year a different office runs a similar evaluation. Companies sometimes mistake this for pipeline. It is not; it is the same opportunity being spent twice, and the second run rarely converts either unless something structural changed. When you recognize this pattern, the useful move is to stop investing in the demonstration and start investing in the procurement path.
It ends in a finding. The pilot revealed something disqualifying: an unexpected outbound dependency, an accessibility failure, an inability to run inside the boundary, a data-handling posture the agency cannot accept. This looks like the worst ending and is often the most useful, because the finding is now specific, dated and fixable, and the same program will usually re-engage once it is fixed. What makes it genuinely bad is only discovering it in the last week when there is no time to respond.
The practical use of this taxonomy is mid-pilot. If you can see by the halfway mark which ending you are heading for, you can change it. Heading for a repeated pilot means the procurement conversation has not happened and there is still time. Heading for a finding means the fix should be scoped now so the response is ready when the report lands.
The mid-pilot conversation that converts
There is a specific moment that decides most of these, and it is roughly two thirds of the way through. Early enough that results are visible and late enough that they are real.
At that point the vendor should bring the sponsor three things in one meeting. The measurement so far against the baseline. The written enterprise scope with a price. And a plain statement of what has to happen, by when, for the enterprise version to be running in the next fiscal year, including the procurement steps and who owns each one.
The purpose is to move the sponsor from evaluating a product to managing a schedule. A sponsor who leaves that meeting with a dated list of actions, some of them theirs, behaves differently from one who is waiting for a final report. And it surfaces the blocking problems while there is still time: that the instrument does not exist, that the budget line is in a different program, that a second approval is needed nobody mentioned.
What the sponsor needs from you two thirds of the way through
Editorial weighting, illustrative rather than measured. The last row is low because new capability at this moment widens the scope and delays the decision.
Pricing the enterprise version so it can be approved
A price that a sponsor cannot defend upward does not convert regardless of how good the pilot was. Three properties make a price defensible.
It maps to something countable. Users, transactions, records processed, sites. A price the program can multiply by a number it already knows is a price it can put in a justification. A negotiated flat number with no visible basis invites a longer review.
It has a smaller first step. The full enterprise scope with an option structure that lets the program start at one division and expand, rather than a single number that requires the largest available approval on the first purchase. Programs buy in the increment their authority covers.
It separates the software from the work. The license or subscription is one thing. Deployment, integration, data migration and training are another, and mixing them makes the recurring cost look larger than it is, which matters in the second year when someone reviews it for renewal.
How we work on this with a software company
Precision Federal builds AI, data platforms, software and cloud systems and delivers them into production, including inside federal agencies. On a pilot that has to convert, we typically do three things.
We design the pilot as a first increment and deliver it. The environment in the production destination or its documented equivalent, federated sign-on, decision and access logging, the outcome measurement computed from real data, a reproducible deployment, and accessibility handled in the components rather than remediated later. Delivered against acceptance criteria written as measurements: the deployment installs from a clean checkout, the metric is computed continuously and matches a manual recomputation, the interface passes an accessibility audit.
We write the security documentation during the pilot. System security plan, control implementation summary, boundary and data-flow diagrams, evidence, and a plan of action for anything not yet met, written against the government control catalog by the engineers who built the deployment. This is the item that most often decides whether the enterprise version starts next quarter or next year.
We write the technical half of the expansion case. What the system looks like at program scale, what it costs to run, what integrations it needs, what the rollout sequence is, and what risks attach to each step. The sponsor needs this to make their argument internally, and it needs to be technically honest because someone in the agency will check it.
What the company keeps. The code, assigned in writing and committed to the company's repositories from day one. The deployment definitions, the documentation package, the measurement code. The agency relationship, entirely; we are visible or invisible in front of the customer as the company prefers. Our pre-existing tooling is named in the agreement, excluded from the assignment, and licensed to the company perpetually so nothing we bring can strand a future maintainer.
How it is priced. Fixed-price milestones fit most pilot work because the scope is definable: the deployment, the measurement, the documentation package, the expansion case. Where the agency relationship is live and the demands are still forming, a committed team for a stated term works better. The first step is one email with a one-page brief: what the product does, which agency and program, what the pilot is supposed to prove, and the date it has to be finished. We return a scoped, priced statement of work.
Bottom line
A federal pilot is the first phase of a purchase or it is a demonstration with a longer timeline, and which one it is gets decided before it starts. Find the sponsor who can buy. Ask the contracting officer what instrument the follow-on rides. Set the metric to something the program already reports and measure the baseline first. Build the pilot where production will live, instrument the outcome, log the decisions and write the security package while you build. Two thirds of the way through, bring the sponsor a result, a priced enterprise scope and a dated list of what has to happen. Pilots designed that way convert. Pilots designed to impress get repeated next year with a different office.
Frequently asked questions
Usually because no purchase path was chosen before the pilot started. The pilot gets funded from a discretionary source, succeeds, and then the follow-on needs an instrument nobody has set up, which takes months. Other common causes: the sponsor was an innovation office with no acquisition budget, the metric was interesting rather than one the program reports upward, the security authorization was deferred until after the pilot, and no enterprise scope or price existed in writing.
Five things. Who the sponsor with a budget is, which may not be who invited you. What instrument the follow-on purchase would ride, asked directly of the contracting officer. What number defines success, chosen from metrics the program already reports, with a baseline measured before you start. Whether the pilot runs in the environment production will use. And a one-page written description of what the system does at full program scale and what it would cost there.
During. Security authorization is the longest step in most federal software purchases, and starting it only after a pilot succeeds adds two or three quarters to the path to production. Writing the system security plan, control implementation summary, boundary diagrams and evidence while the pilot deployment is being built uses that deployment as the evidence. It requires engineers who can write the documents accurately as they build, which is a staffing decision made at the beginning.
Pick a number the program office already reports to someone above it: cycle time on a tracked process, backlog size, error rate against a stated target, cost per transaction. Measure the baseline before the pilot starts, because it is unobtainable afterward. Compute the metric continuously inside the running system from real data, in a way the program can reproduce itself. Accuracy and user satisfaction are inputs; they are not the number that appears in a budget justification.
So the sponsor can defend it upward. Tie the price to something countable the program already knows, such as users, sites, transactions or records, rather than a flat negotiated figure with no visible basis. Structure it so the program can start at the increment its approval authority covers and expand by option. And separate the recurring software cost from one-time deployment, integration and training work, because that separation matters at the first renewal review.
