The prototype worked. The demonstration went well, the sponsor saw the capability, and the transition agreement says a production award can follow. Then the program slows down, and the reason is rarely technical merit. Production is a different contract with a different set of obligations, and the team that proved the capability is usually not the team that can carry it through authorization, integration and sustainment. The prime that plans that handoff before the demonstration gets a production award. The prime that treats production as a continuation of the prototype spends a year discovering what the prototype never had to answer for.
This is written for the program director on the prime's side of that transition. You hold the agreement, the customer relationship and the accountability. The question in front of you is how to divide the next eighteen months of work between your own program organization and the specialist engineering that built the thing, and what evidence has to exist for the follow-on to be justified in writing.
Why prototypes strand
A prototype is scoped to answer one question: can this be done. Everything not needed to answer that question is properly excluded, and a good prototype team excludes ruthlessly. Production asks a different set of questions, and none of them were in scope.
Can it run without the people who built it. Can it be authorized to operate on the network where the mission is. Can it be integrated with the systems that feed it and consume it, on their schedules rather than yours. Can it be sustained for years by an organization with turnover. Can the government defend the sole-source decision that follows the prototype, on the record, if someone asks.
Each of those is a distinct body of work with its own specialists and its own calendar. The failure mode is not that a prime forgets them. It is that the prime assumes the prototype team will do them, the prototype team assumes the prime will, and the first honest conversation happens after the transition window is already narrow.
What most often stops a successful prototype from reaching production
Editorial weighting, illustrative rather than measured. The last row is deliberately low: technical failure is the rarest cause of a stranded prototype.
The split that works
The clean division is not by technology. It is by who owns which obligation, and it holds across programs because the obligations are structural.
The prime owns the government-facing surface. The agreement and its modifications, the schedule the program office believes, the cost and pricing position, the transition documentation, the sustainment organization, the security officer relationship, the integration negotiations with other program offices, and the written case for the follow-on. These are relationship and process obligations. They cannot be subcontracted in any meaningful sense, because the government is buying the prime's accountability for them.
The specialist owns the engineering surface. The code and its tests, the model or algorithm and its evaluation suite, the data pipeline and its lineage, the deployment automation, the control implementation evidence, the interface adapters, the performance work, and the technical content of every artifact the prime hands the government. These are depth obligations. A program organization staffed for oversight cannot produce them at the pace production requires without either hiring the specialists or partnering with them.
The line between the two is a document flow, not a fence. The specialist writes the technical content; the prime shapes it, owns it, and signs it. That distinction matters because it is what lets a prime carry a specialist on a program of record without diluting its own accountability.
What actually changes from prototype to production
Naming the deltas concretely is the most useful thing a program director can do in the first month of transition planning. Six recur on nearly every program.
Authorization stops being optional. A prototype often runs in a development enclave under an interim posture, or in a lab that never touched operational data. Production runs where the mission is, and that means a security control baseline, a system security plan, a control implementation record, continuous monitoring, and a security officer who signs. If the prototype was built without knowing which baseline and which authorizing environment it was headed for, the first six months of production are re-engineering rather than new capability.
Interfaces stop being mocks. Prototype teams stub the systems they do not control, and they should. Production means real interface agreements with real program offices that have their own release trains, their own change boards, and no obligation to move on your schedule. Every mocked interface is a scheduling dependency nobody has priced.
Data stops being a snapshot. Almost every prototype runs against an extract someone pulled once. Production needs a live feed, with a schema contract, a backfill story, monitoring on freshness and volume, and a defined behavior when the upstream system changes without telling you. The engineering under a live feed is usually larger than the engineering under the capability it feeds.
Failure stops being interesting and starts being an incident. A prototype that falls over gets restarted by the person who built it. Production needs health checks, alerting with a human on the other end, runbooks written for someone who was not there, defined degraded modes, and a recovery time the program office has agreed to.
Change stops being free. Prototype code changes daily. Production code changes through a pipeline with tests, review, a release record, and a rollback. That discipline costs velocity and buys the ability to hand the system to somebody else, which is the entire point of production.
People stop being irreplaceable. A prototype can live in three people's heads. A production system that lives in three people's heads is a risk finding. Documentation, test coverage and deployment automation are what convert a capability into an asset the government owns.
Where the production effort actually goes, relative to the prototype build
Editorial weighting, illustrative rather than measured. The last row is deliberately low: production adds discipline, not features.
The evidence the prototype has to produce
Under the other transaction authority framework, a follow-on production award can proceed without further competition when the prototype project was competitively awarded, all participants were notified that a follow-on could result, and the prototype project was completed successfully. The third condition is where programs get into trouble, because "successfully completed" is not self-evident and somebody in the contracting shop has to write it down and defend it.
The practical answer is to design the prototype's evidence output as a deliverable from the start, on the theory that the follow-on justification is a document your team is going to write and should therefore plan. What that means in practice:
- Written success criteria, agreed before the work starts. Not "demonstrate feasibility." A named threshold on a named measure, evaluated on a named dataset or in a named environment, with the method of measurement stated. If the agreement's statement of work carries adjectives, propose numbers and get them incorporated by modification.
- A test record, not a demonstration memory. The demonstration is theater in the good sense, and it matters. It is not evidence. Evidence is a repeatable evaluation run with logged inputs, versioned code, and results that someone else can reproduce from the artifacts you delivered.
- Performance measured where it will run. Latency and throughput on the target infrastructure with representative load, not on a workstation with a favorable dataset. A number from the wrong environment invites a question you cannot answer.
- An honest limitations section. What the prototype did not test, and what the production phase will therefore have to prove. This makes the follow-on justification stronger, not weaker, because it shows the program office understood what it bought.
- The prototype's own artifacts as a starting baseline. Architecture description, data flows, interface list with status, control implementation notes, and the backlog of everything deliberately deferred. That deferred list is the production scope, already written.
Everything on that list is technical content. That is why the specialist has to be in the room for the transition planning, not brought back afterward to fill in a template.
Two ways to structure the production team
Once the transition is real, a prime has a decision to make about how the specialist engineering sits on the production contract. Both shapes are workable and they fail differently.
| Dimension | Specialist as scoped subcontractor | Specialist as committed embedded team |
|---|---|---|
| What is bought | Defined outcomes with acceptance criteria, priced to milestones | Named engineers at a committed allocation on the prime's plan |
| Best when | Scope is knowable: hardening, authorization evidence, a named integration | Scope moves with the program office and the backlog is negotiated monthly |
| Schedule risk | Held by the specialist inside each milestone | Held by the prime, who is setting priorities |
| Estimating burden | Higher up front; the scope has to be written before pricing | Lower up front; the arguments move to sprint planning |
| Typical failure mode | Scope written in adjectives, so acceptance never arrives | Bought as capacity, then held accountable for an outcome nobody scoped |
| Fits the evaluation | Reads as a real workshare with an evaluated technical scope | Reads as key personnel and a staffing plan |
Most production programs use both. Authorization evidence, a specific integration, and the deployment automation are milestone work with clean acceptance criteria. Sustainment engineering and the negotiated backlog are a committed team. Writing them as two instruments, or as two clearly separated sections of one, prevents the most common argument on these programs, which is whether a given piece of work was already paid for.
The authorization work, described honestly
This is the part program directors most often underestimate, so it is worth being specific about what the work actually is.
A production system needs an authorization decision from an official who will personally accept the residual risk. To get one, the system needs a categorization, a selected control baseline with tailoring, an implementation of each control that is either technical, inherited from the hosting platform, or covered by a procedure, evidence that each implementation exists, an assessment by someone independent of the build, a plan for anything not fully satisfied, and a monitoring approach that keeps the picture current after the decision. NIST SP 800-53 supplies the control catalog most federal environments use.
Two facts govern the schedule. First, most controls are satisfied by the hosting platform rather than by your code, so choosing a platform with an existing authorization removes a large fraction of the work before you write a line. Inheritance is the single largest lever on an authorization timeline. Second, the evidence is produced by engineering, not by writing. A control that says logs are retained and reviewed is satisfied by a logging pipeline, a retention configuration and a review procedure someone actually follows. Writing that a control is implemented, when it is not, produces a finding and a delay.
For an AI or machine-learning component there is a third layer. The authorizing official will ask what happens when the model is wrong, how you know it is still performing, and who reviews its outputs. That is answered with an automated evaluation suite that runs on a schedule, drift monitoring on inputs and outputs, a documented human review step for consequential decisions, and a rollback to a previous model version that has been rehearsed. Building those during the prototype is inexpensive. Retrofitting them into a production system under an authorization deadline is not.
How we work inside a prime's program
Precision Federal builds AI, data platforms, software and cloud systems and delivers them into production inside federal agencies. On a prototype-to-production transition we come in as the specialist engineering under the prime, and the shape is deliberately plain.
The first three weeks. We read the agreement, the prototype's artifacts and whatever the program office has said about the operational destination. We produce three things: a written gap list from prototype to production organized by obligation rather than by component, a proposed authorization path naming the hosting environment and the inheritance we expect from it, and an interface register listing every external system with its current status, its owner, and the lead time to a real agreement. That document is what makes the transition plan estimable, and it is useful to the prime whether or not we do the build.
What we deliver after that. Hardened code with test coverage, deployment automation that runs from a clean checkout, the data pipeline with monitoring and a schema contract, the evaluation suite and its scheduled runs, the control implementation evidence in the form the assessor wants, the interface adapters, and the runbooks. We write the technical content of the CDRLs; the prime shapes and delivers them.
What the prime keeps. The customer relationship and every conversation with the program office that the prime wants to own. The code, delivered with a written assignment rather than a recital, with our pre-existing tooling named, carved out, and licensed to the government for use in the delivered system so no future maintainer is blocked. The data, which was never ours. The right to have someone else maintain the system, which is what the documentation and the deployment automation exist to make real.
How it is priced. Fixed-price milestones where the scope is written, with acceptance criteria stated as measurements rather than adjectives. A committed team at a named allocation where the backlog is negotiated. We do not price a defined outcome as hours and then discover the accountability question at acceptance.
How to start. One email with a one-page brief: what the prototype proved, where the production system has to live, what it has to talk to, the authorization destination if one has been chosen, the date that matters, and who can approve a scope change. We come back with a scoped, priced statement of work. If the answer is that the transition is not ready to be scoped yet, we say that too, and say what would make it ready.
Five failure modes worth naming
The demonstration becomes the plan. A successful demonstration creates confidence that the hard part is done. It is the easiest part. Treating the demonstration as the midpoint rather than the finish line is the single most useful correction a program director can make.
The prototype team is released before the transition plan exists. The knowledge that makes the gap list accurate lives with the people who built the thing. Releasing them to save burn before that document is written costs more than it saves, every time.
The authorization is treated as paperwork. Assigning it to a compliance function with no engineering attached produces a document set that does not match the system, which is discovered at assessment. Authorization evidence is an engineering deliverable with a writing component, not the reverse.
Sustainment is left as a line item with nobody in it. A production system needs a named organization with funded people, on-call coverage, a patching cadence and a plan for the third year. If the plan does not say who, the government will ask, and the answer shapes the follow-on decision.
Data rights are settled at the end. What the government receives, in what form, with what license, and what remains the contractor's, is cheap to agree in the transition plan and expensive to argue during closeout. Name the pre-existing tooling early and license it clearly.
What the transition plan should contain
A transition plan that a program office can act on is short and specific. It names the operational environment and the authorizing official's organization. It lists every interface with an owner and a status. It states the sustainment organization and its funding line. It carries the gap list with an estimate against each item. It states the success evidence from the prototype and where it lives. It names the key people on both the prime and the specialist side with their allocations. And it says what the government owns at the end, in the words of the delivery clause rather than in a summary.
What to leave out: a restatement of the prototype's technical merit, which nobody is disputing, and a schedule with no dependency structure, which nobody believes.
Bottom line
Production is not more prototype. It is a different set of obligations, and the ones that strand programs are authorization, integration, sustainment and evidence rather than capability. The division that works puts the government-facing surface with the prime and the engineering surface with the specialist, and connects them with a document flow the prime owns and signs. Design the prototype's evidence output as a deliverable so the follow-on justification writes itself. Choose the hosting environment for what its authorization lets you inherit. Name the sustainment organization before the program office asks. A prime that does those four things during the prototype, rather than after it, has a production award to win rather than a transition to survive.
Frequently asked questions
Under the other transaction authority framework, a follow-on production award can proceed without further competition when the prototype project was competitively awarded, all participants were notified that a follow-on production award could result, and the prototype project was completed successfully. The third condition is the one that requires planning, because someone has to write down what successful completion means and support it with evidence. Agree the success criteria as measurable thresholds before the prototype starts, and produce a reproducible test record rather than a demonstration memory.
Rarely because the capability did not work. The usual causes are structural: no authorization path was chosen while the prototype was being built, so production begins with re-engineering; interfaces to fielded systems were mocked rather than negotiated, so integration is an unpriced dependency; no sustainment organization was named or funded; and the prototype's evidence does not support a written follow-on justification. Each of those is preventable during the prototype at a fraction of the cost of fixing it afterward.
By obligation rather than by technology. The prime owns the government-facing surface: the agreement, the schedule, the cost position, the sustainment organization, the program office relationship, and the written case for the follow-on. The specialist owns the engineering surface: code and tests, the evaluation suite, the data pipeline, deployment automation, control implementation evidence, interface adapters, and the technical content of every deliverable. The prime shapes, owns and signs what the specialist writes.
The standard package plus three additions. The standard package is categorization, a tailored control baseline from a catalog such as NIST SP 800-53, an implementation of each control that is technical, inherited or procedural, evidence for each, an independent assessment, a plan for anything unsatisfied, and continuous monitoring. For a model, the authorizing official will also ask what happens when it is wrong, how you know it is still performing, and who reviews consequential outputs. That means a scheduled evaluation suite, drift monitoring, a documented human review step, and a rehearsed rollback to a prior model version.
Usually both, written as separate scopes. Work whose scope is knowable, such as hardening, authorization evidence and a named integration, fits fixed-price milestones with measurable acceptance criteria, and the schedule risk sits with the specialist. Work whose backlog is negotiated with the program office month to month fits a committed team at a named allocation, and the schedule risk sits with the prime who is setting priorities. Separating them prevents the most common dispute on these programs, which is whether a piece of work was already paid for.
