Two years into a large AI program, an honest inventory usually reads like this. A governance framework exists and is approved. A council meets. Several thousand people have been trained. There is a portfolio of pilots, most technically successful, a few genuinely impressive. And when the chief financial officer asks which line in the accounts has moved, the answer takes a long time and does not fully land. The program has produced capability, awareness and readiness. It has not produced systems the business runs on, which is the only form in which any of this becomes money.
This is not a failure of ambition or of talent. It is what happens when a program is architected around enablement instead of around shipping. The change that fixes it is structural and it can be made at any phase review. Precision Federal builds the software inside programs like these. What follows is the program architecture we see work, phase by phase, and what a sponsor should demand at each review.
Why pilots accumulate and systems do not
A pilot is easy to start and easy to end. It needs a willing business owner, a modest budget, a data extract and a few months. It ends with a demonstration and a finding, and whether it goes further is somebody else's decision. Nothing in that structure produces a running system, and nothing in it is anyone's fault.
A production system needs a different set of things, and each of them is a queue with an owner outside the program. It needs somewhere to run that operations will support. It needs to pass security review and, in regulated or public sector settings, to reach an authorization. It needs identity, monitoring, on-call, backup and a deployment pipeline. It needs the upstream data to arrive reliably rather than as a one-time extract, which usually means changing a system somebody else owns. It needs a business process to change, and people to be trained on the new one. It needs a budget line for running it, not just for building it.
None of those are in a pilot's scope, and a program that funds pilots one at a time funds none of them. The gap is not technical difficulty. It is that the twelve things standing between a working prototype and a running system are each owned by a different part of the organization, and nobody's job is to clear all twelve for a specific system by a specific date.
That observation determines the program architecture. Somebody's job has to be exactly that.
What predicts whether a program ends with software the business runs on
Editorial weighting, illustrative rather than measured. The last row is the metric programs report most often and the one that predicts shipped software least.
The four parts of a program that ships
A program that ends with running software has four components, and the common failure is having three.
A platform team that owns the path to production. Permanent, small, and staffed by engineers rather than coordinators. Its product is not a system; its product is the ability to ship systems. That means the environments, the deployment pipelines, the identity and access patterns, the data access layer, the monitoring, the automated evaluation suite for models and prompts, and the pre-agreed security pattern that a new system inherits instead of negotiating from scratch. Each use case that lands after the first should be materially faster because this team exists. If the fifth system takes as long as the first, the platform team is not doing its job and the program is really five projects wearing a program's clothes.
A small number of production use cases. Not a portfolio. Three to five in the first eighteen months, chosen deliberately, each with a business owner who will have to use the result. The discipline of fewness is the hardest part of this to sell politically and the most important. Ten pilots and three shipped systems is a worse program than four use cases all of which ship, because the ten pilots consumed the same money and left nothing running.
An engineering partner accountable for shipping. Whether internal or external, someone must be accountable for named systems reaching production by named dates against written acceptance criteria. Accountability means specific: this system, these criteria, this date. A partner accountable for advice is not accountable for shipping, and a team accountable for a workstream is not accountable for a system.
A benefits ledger the finance function owns. Not a benefits case written once to secure funding. A live ledger, maintained with the finance team, that names for each system the measure, the baseline as it stood before the system existed, the current value, the method, and who signs off that the value is real. This is the component most often missing entirely, and its absence is why programs cannot answer the chief financial officer's question.
Choosing use cases for measurability, not for enthusiasm
Selection is usually run as a prioritization exercise scoring value against feasibility, which produces a list ranked by optimism. A better filter is short and it is mostly about whether anyone will be able to tell if it worked.
Five tests, applied honestly, cut a long list quickly.
- Is there a named person whose work changes? Not a department, a person in a role who will use the output and can say whether it helped. If the benefit accrues to nobody in particular, nobody will adopt it and nobody can measure it.
- Can the benefit be measured with a number that already exists? Handling time, error rate, throughput, cycle time, cost per unit. A benefit requiring a new measurement system is a benefit that will be argued about for a year. Where the baseline is not currently captured, measuring it for six weeks before the build starts is cheap and it is the difference between a defensible claim and a debate.
- Does the required data exist, arrive reliably, and may it be used for this? All three. Data that exists but arrives by manual extract is a system with a person inside it. Data whose permitted use is unclear is a governance review sitting between the build and production.
- Is the accuracy the business needs achievable and known? If the decision requires near-certainty and the achievable performance is unknown, that is a pilot question and should be answered in a short, cheap phase before the build is funded.
- Will the work actually change, and who will make it change? The most common cause of a shipped system with no measured benefit is that the old process continued alongside the new one. Somebody senior in the business has to retire the old path, and that person should be named at selection, not found later.
Use cases passing all five are usually unglamorous. They are also the ones that produce numbers the finance function will sign, and a program with two signed numbers in year one has bought itself the credibility to attempt harder things in year two.
Phasing, and what to demand at each review
The phases below are a shape, not a calendar; the durations depend on the size of the organization and the state of its data. What matters is that each review has a demand attached to it that cannot be satisfied with a document.
| Phase | What is built | What the review must see | Stop condition |
|---|---|---|---|
| Foundation | Platform skeleton, security pattern, first use case selected and specified | A trivial service deployed through the real pipeline into the real environment | No agreed security pattern, or no named business owner |
| First system | One production system, end to end, with its benefit baseline measured | Real users using it in their actual work, and the baseline recorded | Adoption below the level at which the benefit exists |
| Repeatability | Two or three more systems, reusing the platform | Each landing measurably faster than the first, with the reuse itemized | Second system took as long as the first |
| Scale | Wider rollout, self-service platform capabilities, internal team growth | Systems shipped by teams outside the original program | Everything still routes through the same few people |
| Steady state | Run, support and improvement inside the business | A funded run budget and a named owner per system | No operating budget, meaning the systems will decay |
Three demands are worth making at every review regardless of phase.
Show me the system, not the slide. Running software, in the real environment, used by a real user doing real work. This one demand changes program behavior more than any governance instrument, because it cannot be satisfied by preparation.
Show me the ledger. For each shipped system: the measure, the baseline, the current value, the method, and the finance signature. If the ledger is a forecast rather than a measurement, say so out loud at the review.
Show me what the platform made cheaper. Specifically: which components the newest system inherited rather than built. If the answer is vague, the platform is a name for a team rather than an asset.
The benefits ledger the finance function will believe
Benefit claims from technology programs are treated skeptically by finance functions for good reason: they are usually built on a counterfactual nobody can test, and they usually count time saved as money without anyone's cost actually falling. Building a ledger the chief financial officer will sign is not hard, but it has to be agreed before the build rather than assembled afterward.
Four elements make a claim credible.
A baseline measured before the system existed. Not reconstructed from memory or from a vendor's benchmark. If the metric is handling time, measure handling time for several weeks first. This costs almost nothing and it is the single item that decides whether the claim survives scrutiny.
A method agreed with finance in advance. How the measurement is taken, over what period, with what adjustment for seasonality or volume changes, and what counts as attributable to the system rather than to other changes happening at the same time. Agreeing the method beforehand removes the incentive to choose the flattering method afterward.
Honesty about the form of the benefit. Capacity released is not cost removed unless the organization actually reduces cost or redirects the capacity to something it would otherwise have hired for. Say which one is claimed. A ledger that distinguishes cash saved, cost avoided, revenue enabled and capacity released is far more persuasive than one that adds them together into a single number, because the person reading it knows the difference and will discount the whole thing if the distinction is missing.
Run cost on the same page. Infrastructure, model or service usage, support, monitoring and the ongoing engineering to keep the system correct as the data underneath it changes. A benefit claim with no run cost next to it reads as advocacy. Including it, and still showing a positive number, reads as analysis.
Set a review point per system, a fixed number of months after go-live, at which the measured value is recorded whatever it says. Programs that publish a system that underperformed, with the reason, get more credibility than programs that publish only wins, and the credibility is what funds the next phase.
Why systems stall between a working prototype and production use
Editorial weighting, illustrative rather than measured. The last row is low because model quality is rarely what stops a system that already worked in a pilot.
Governance that helps rather than delays
Governance is necessary and it is usually built in a shape that slows shipping without improving safety. Two changes make it useful.
First, make governance a pattern rather than a review. A system that uses the approved identity model, the approved data access layer, the approved logging and the approved evaluation suite should be able to reach production without a bespoke security negotiation, because the pattern was reviewed once. Systems that deviate get scrutiny proportional to the deviation. This is the single highest-return investment the platform team makes, and it converts governance from a queue into an inheritance.
Second, tier the review by consequence. A system that drafts internal summaries and a system that influences a decision affecting a customer's money or rights are not the same risk and should not face the same process. Where model risk management practice applies, as in banking, the tiering is already familiar and the program should map onto the existing framework rather than inventing a parallel one. In federal settings the control obligations and the authorization path are known in advance and belong in the platform pattern from the first system, not discovered at the end of the first build.
What governance should insist on for every production system is short: a named accountable owner, a written statement of what the system does and does not decide, an automated evaluation suite that runs on a schedule rather than once, monitoring that detects drift in the inputs as well as the outputs, a documented path for a human to override or escalate, and a logged record sufficient to explain a specific decision months later. Those six are checkable and they are the ones that matter when something goes wrong.
How an engineering partner fits into the program
Precision Federal builds AI systems, data platforms, cloud infrastructure and full-stack web and mobile software, and delivers them into production, including inside U.S. federal agencies where a system must pass security authorization, handle controlled information and meet accessibility requirements. In a transformation program that experience matters in a specific way: we treat the deployment destination, the control set and the evidence obligations as first-phase design inputs, because in the environments we work in there is no other option, and the same discipline is what stops a commercial pilot from dying in security review.
Inside a program we take one of two roles, sometimes both. We build the platform: the environments, pipelines, identity and data access patterns, the evaluation suite and the security pattern a new system inherits, so that the third system lands faster than the first. Or we are accountable for named production systems, delivered in fixed increments against acceptance criteria written as tests, billed at milestones tied to demonstrable events, with security and accessibility work inside the increments rather than deferred to a final review.
In the first weeks of an engagement there is a trivial service deployed through the real pipeline into the real environment, which retires most of the assumptions everyone is carrying about what production actually requires. Within about a month there is a working slice of a real use case running against real data, shown to the people who will use it.
Your organization keeps everything. You keep the code and the intellectual property, transferred by a written present assignment rather than a work-for-hire recital, with our pre-existing tooling named, carved out and licensed to you perpetually. You keep the data, held inside the boundary we agree and never used to train anything. You keep the client relationship where we work behind an advisory firm, and we do not approach that client independently without your agreement. Handover is a rehearsal in which your team deploys while we watch, before the final milestone is paid.
The first step is one email with a one-page brief: the use case, the systems and data involved, the deployment destination and any authorization the result must reach, the date that matters, and who can approve a change. We return a scoped, priced statement of work with acceptance criteria written as tests. No call required.
Bottom line
Programs produce what their architecture is shaped to produce. An architecture of training, councils and a pilot portfolio produces training, councils and pilots, all of which are real and none of which is a system the business runs on. Four components change the output: a platform team whose product is the ability to ship, a deliberately small number of use cases chosen because the benefit is measurable and someone's work actually changes, a party accountable for named systems reaching production against written criteria, and a benefits ledger the finance function agreed to before the build. Demand running software at every review, a measured ledger rather than a forecast, and a specific account of what the platform made cheaper. Those three demands, asked consistently, reshape a program faster than any restructure.
Frequently asked questions
Because a pilot needs a willing owner, a data extract and a few months, while a production system needs somewhere to run that operations supports, a passed security review, reliable upstream data, identity and monitoring, a changed business process and a funded run budget. Each of those is owned by a different part of the organization, and in most programs nobody's job is to clear all of them for a specific system by a specific date. Making that somebody's explicit accountability is the structural fix.
Select for measurability rather than enthusiasm. Ask whether a named person's work changes, whether the benefit can be measured with a number that already exists, whether the data exists and arrives reliably and may lawfully be used this way, whether the accuracy the decision requires is known to be achievable, and who will retire the old process. Use cases passing all five are usually unglamorous and are the ones that produce numbers a finance function will sign.
Three things, at every review. Running software in the real environment used by a real user doing real work, rather than a demonstration or a slide. The benefits ledger showing measure, baseline, current value, method and the finance sign-off, with any forecast identified as a forecast. And a specific account of which components the newest system inherited from the platform rather than built, since a vague answer means the platform is a team name rather than an asset.
Measure the baseline before the system exists rather than reconstructing it afterward, agree the measurement method with the finance function in advance so the flattering method cannot be chosen later, distinguish cash saved from cost avoided from revenue enabled from capacity released instead of summing them, and put run cost on the same page. Then set a review point a fixed number of months after go-live and record the measured value whatever it says, including when it disappoints.
The ability to ship, not any single system. Concretely: the environments, deployment pipelines, identity and access patterns, the data access layer, monitoring, the automated evaluation suite for models and prompts, and a pre-agreed security pattern a new system inherits instead of negotiating from scratch. The test of whether it is working is whether each system lands materially faster than the one before. If the fifth takes as long as the first, the program is several projects rather than a program.
