The gap is almost never the engineering
A commercial product that has survived a few enterprise procurements is usually further along the federal bar than its own team believes. Access is federated through an identity provider with MFA. Data is encrypted in transit and at rest. Changes go through pull request review and a pipeline that runs tests. Logs land somewhere queryable and stay there for a year. Vulnerabilities get patched on a rhythm the team can describe from memory. Those are real implementations of real controls. What the team does not have is a document that says so in the sentence structure a federal reviewer is trained to read, with a pointer to the artifact that proves it.
That gap costs real money and real calendar time, in a specific way. Building the genuinely missing pieces is typically 10 to 20 percent of the effort. The rest is describing what already exists, finding the config export or the ticket that evidences it, and reconciling the four names your organization uses for the same system. Teams price this backwards. They budget a six-month build and discover after two months that they are running a documentation project with a small engineering tail.
It feels like an engineering problem because the request arrives in engineering language. A program office asks for a system security plan aligned to NIST SP 800-53 Rev. 5, or 800-171 Rev. 3 if the data is CUI on your systems. A civilian agency points at the NIST AI Risk Management Framework. A bank-adjacent buyer wants something that looks like SR 11-7 model validation. A DoD program expects an SBOM in CycloneDX or SPDX. Each request names a technical standard, so an engineering leader reasonably reads it as a build. It is a build in maybe one case out of five.
Where the work actually sits on a typical commercial-to-federal readiness effort
Editorial weighting from practitioner reading of control-family assessments. Illustrative of the shape, not a measured statistic across a sample of firms.
What a reviewer is actually reading for
A control narrative is not prose about your architecture. It answers four questions and stops. Who or what performs the control. How, mechanically, naming the systems that do it. How often, or on what trigger. Where the proof lives. A reviewer scanning a 300-control package spends under two minutes per control on a first pass. If those answers are not findable in that time, the control gets a question, and questions are the currency of delay.
A weak narrative reads: "The system enforces least privilege through role-based access control." That could be true of anything, including a product with one role called admin. A usable one reads: "Access is provisioned in the identity provider against six roles mapped in the role definition table. Production database credentials are issued only to the two roles marked production-eligible. The engineering manager reviews the full role assignment export quarterly and signs it. The signed export is stored in the compliance repository at the path named in the evidence index." That answers all four questions and points at a file. It takes a reviewer twenty seconds.
Teams that write the first kind of sentence are not being lazy. They are writing for a reader they imagine wants to understand the system. The reviewer is deciding whether to accept residual risk on behalf of an agency, and writing down why.
The evidence problem is worse than the writing problem
Most teams underestimate this part. A narrative takes a day per control family from someone who knows the system. Evidence takes longer, because it has to be a specific thing captured at a specific time, and most of what a running engineering organization produces is ephemeral.
Your access reviews happen in a Slack thread. Your incident response ran last March and the retro lives in a page somebody has since renamed. Your penetration test was scoped to the marketing site. Your vulnerability scanner reports into a dashboard that shows current state and cannot show what state it was in six weeks ago, which is exactly what an assessor wants, because current state can be arranged. Your configuration baseline lives in Terraform, which is good, but nobody has exported the applied state as a dated artifact.
The fix is mechanical. Pick a date. From it forward, every recurring control writes a dated artifact into one repository with a stable path, and the evidence index maps each control identifier to that path. A quarterly access review becomes a signed CSV. A vulnerability cycle becomes a dated scan export plus the remediation ticket list. A change becomes the pull request URL plus the approval record. None of this changes what your team does. It changes whether what your team does leaves a trace that survives the calendar.
Budget six to ten weeks for the first evidence cycle, because you cannot fabricate a quarterly review that has not happened. This is the most common reason a readiness effort slips: the writing finishes in week five and the package still cannot ship until an evidence period has elapsed. Start the clock on day one, in parallel with the writing.
The AI layer has the same shape and a shorter history
If the product has a model in it, a second documentation stack applies, younger and less settled than the security one. The NIST AI Risk Management Framework is the common reference. Two of its four functions, Govern and Map, are almost entirely documentation. Govern asks who owns the risk, who approves deployment, and what the exception path is. Map asks what the system is, who uses it, what data feeds it, and what decisions it touches. A team shipping a model for three years can usually answer all of that in a meeting and has never written it into one document.
ISO/IEC 42001 covers the same ground as a certifiable management system, and some buyers now ask for it by name. If the model is credit-adjacent or risk-scoring, SR 11-7 is the older and more demanding lineage, and it wants three things federal AI documentation often skips: conceptual soundness of the approach, ongoing monitoring with defined triggers, and outcomes analysis comparing predictions to what actually happened. For anything with a language model in the loop, the OWASP Top 10 for LLM Applications and MITRE ATLAS give a threat taxonomy reviewers recognize. "We tested it for safety" is not a finding a reviewer can accept. "We ran the ATLAS technique list against the deployed configuration, and here are the results" is.
The thing that trips commercial AI teams is subgroup performance. You have aggregate accuracy, and you have it by customer segment because the business asked for that. You almost certainly do not have it broken out along the dimensions a federal reviewer asks about, and you may not have the labels to compute it. That is one of the few places the answer really is engineering work, and it is worth finding out early rather than in week eleven.
Four names for the same system
Every organization past about 80 people has this problem and none of them know it until someone tries to write a system boundary. The product is called one thing in the marketing material, another in the cloud account tags, a third in the on-call rotation, and a fourth in the contract. The database holding customer records has a legacy name from an acquisition. Two teams disagree about whether the reporting service is inside the product or a separate internal tool.
None of that matters until you draw a boundary, and then all of it matters at once, because the boundary decides scope and scope decides cost. A boundary that sweeps in the corporate identity provider, the ticketing system and the observability stack is three times the assessment of one that does not. That drawing is a design decision with a price attached, worth a week of a senior architect's time before anyone writes a control narrative. Our piece on drawing a realistic CUI boundary works through the same logic on the regulated-data side.
The deliverable that resolves it is a one-page system description with a component inventory: every component named once, its canonical name, its aliases, whether it is inside the boundary, and who owns it. Circulate it and watch the arguments surface. Those arguments are the actual work. Once the names are settled, the writing goes four times faster because every narrative can refer to components by a name that means one thing.
Typical elapsed weeks by workstream on a first federal documentation package
Workstreams overlap. Total elapsed time is usually 12 to 16 weeks, not the sum, and the evidence cycle is the item that cannot be compressed by adding people.
What this costs, honestly
For a single product with a clean boundary and a cooperative engineering team, a first package aligned to 800-171 or a moderate 800-53 profile runs 400 to 900 hours of specialist time. At market rates for people who have done this before, that is roughly $80,000 to $220,000. Add 150 to 300 hours for a model needing AI RMF or 42001 documentation, and more if the boundary is contested or the product spans cloud accounts owned by different teams.
Full FedRAMP authorization is a different order of magnitude. Public estimates for a moderate baseline run well into seven figures across preparation, third-party assessment and the first year of continuous monitoring, with elapsed time measured in quarters. Most companies asking about federal documentation do not need that first. They need a package good enough for an agency authorization on a specific program, which is smaller and reuses most of the same material if FedRAMP later becomes the path.
The number that should worry a budget owner is not the specialist cost. It is the internal engineering hours consumed answering questions. A poorly run effort pulls 20 to 40 hours a week from senior engineers for three months, because the writer has to ask about everything. A well-run one front-loads a two-day architecture session, reads most of what it needs from configuration and code, and holds engineering time under eight hours a week. That difference is bigger than the entire consulting fee.
Read the code, not the wiki
The practice that separates a fast effort from a slow one is whether the people writing can read the system directly. Terraform and CloudFormation give the network topology, encryption settings, retention policies and IAM structure without anyone being interviewed. CI configuration says what runs on every change. Identity provider exports give the role structure. Dependency manifests generate the SBOM. Access logs say whether the review policy in the wiki is what actually happens.
A team that cannot read those artifacts has to interview its way to every fact, and interviews produce the aspirational version of the system rather than the deployed one. That is where findings come from later. The wiki says quarterly access reviews. The IAM data says the last one was fourteen months ago. An assessor finds that in an afternoon, and it is worse than having no narrative at all, because now the document is wrong.
This is also the honest test to apply to anyone you hire for this work. Ask them what they will read first. If the answer is "your existing policies," the effort will be slow and the output will describe an organization that does not exist. If the answer names your infrastructure repository, your identity provider and your pipeline configuration, they have done this before.
Write for the reuse, not just the submission
A package written for one solicitation gets thrown away. A package written as a maintained asset serves the next five. The structural choices that make the difference are small and cost nothing at authoring time.
Keep the narratives in source control as text, not in a Word document that lives in a shared drive. Keep control text separate from the framework mapping, so the same paragraph about how encryption works can be mapped to 800-53 SC-13, 800-171 3.13.11, and an ISO 27001 annex control without being rewritten three times. Keep the evidence index as a data file rather than a table pasted into prose. Tag each narrative with the date it was last verified against the system, because the question a reviewer asks on year two is not "what does it say" but "when was this last true."
Done that way, the next agency's request is a mapping exercise of a week or two. Done the other way, the second package costs almost as much as the first, and by the third somebody proposes hiring a compliance team, which is how a documentation problem becomes a permanent headcount line.
What to do in the first two weeks
Name the boundary and get the component inventory circulated, argued over, and settled. Start the evidence clock on every recurring control so the cycle is running while the writing happens. Pull the configuration artifacts into one place so the people writing can read the deployed system rather than the described one. Identify the small number of controls that are genuinely missing, which for most products is a handful, and put them in a plan with dates. And find out early whether the model needs subgroup metrics you cannot currently compute, because that is the item with the longest lead time and the one most likely to surprise you.
Then write. The writing is the visible part and the least of it.
Bottom line
Most commercial products that stall at a federal door are not underbuilt. They are underwritten. The controls are running, the reviews are happening, the tests are passing, and none of it exists in a form that lets a reviewer accept risk and sign. Treating that as an engineering problem produces a six-month build for something that was 80 percent documentation. Treating it as a documentation problem with a small engineering tail, run by people who read infrastructure code instead of interviewing their way to the answer, gets a package in twelve to sixteen weeks that the next program can reuse.
Frequently asked questions
Not always. FedRAMP applies to cloud services the government consumes as a service. Software delivered into an agency-operated environment, or work performed under a contract where the agency owns the system, is typically authorized through the agency's own process against 800-53. Ask which path the program is on before budgeting, because the two differ by an order of magnitude in cost.
Twelve to sixteen weeks of elapsed time for a single product with a clean boundary. The compressible parts are the writing and the engineering. The part that is not compressible is one full cycle of recurring evidence, which is why the evidence clock should start on day one rather than after the narratives are drafted.
It counts as strong signal that the controls exist, and the questionnaires you answered are useful raw material. It does not substitute for the package, because federal reviewers work from a specific control catalog and expect narratives tied to specific control identifiers with evidence pointers. The good news is that the underlying facts are almost always already true.
Add a second stack of documentation on top of the security package: intended use, training data provenance, evaluation methodology, performance including subgroup breakdowns, known failure modes, human oversight, and monitoring. NIST AI RMF is the common reference, ISO/IEC 42001 if a buyer asks for it by name, and SR 11-7 if the model informs credit or risk decisions. Budget an extra four weeks, more if subgroup metrics need to be built.
They can, and the facts they hold are irreplaceable. What usually goes wrong is that the work competes with a product roadmap and loses, so a twelve-week effort becomes a nine-month one. The efficient split is a specialist who reads the deployed system and drafts, with your engineers reviewing in bounded sessions rather than being interviewed continuously.