Skip to main content
Business Development

Buying AI consulting: a scope template that prevents rework

You are buying a result you cannot inspect before it exists. The scope of work is the only instrument that fixes that. Here is what goes in it, section by section, and what each omission costs later.

Why AI scopes go wrong in a particular way

With a roof replacement or a network refresh, a buyer who knows nothing about the trade can still recognize a finished job. With an AI system the deliverable is a number on a slide, produced by the people you are paying. Almost every dispute in an AI engagement traces back to a scope that never said what number would be produced, on what data, judged by whom. The template below is the fix, and it is boring on purpose.

A normal software scope describes functions: build these screens, expose this API, hold this uptime. The buyer reads the list and pictures the finished thing. An AI scope has an unknown in the middle of it: whether the records the organization holds will support the result the organization wants. Nobody knows that when the scope is written, the vendor included. A scope written as though the answer were known gets renegotiated later, after the money is spent.

Three unknowns cause most of the trouble. Data condition: how complete the records are, how consistently they were labeled, whether the fields that matter are populated in the years that matter. The achievable ceiling: how well any method could do here, a different question from how a vendor's method scores on a public benchmark. And the integration surface: where the output lands, who approves it, which system has to accept it.

Rework risk by scope section when left vague

Objective with no measurement method
94%
Data access with no date or named owner
90%
Acceptance criteria absent or subjective
86%
Security level settled after work starts
79%
Data rights left to the default clause
74%
Changes handled by verbal direction
68%

Editorial weighting from public acquisition guidance and practitioner reading, not a measured statistic.

Write an objective you can measure

FAR subpart 37.6 directs agencies to describe service work in terms of required results and measurable performance standards rather than methods. That is easy to agree with and hard to follow for AI, because the natural way to write an AI objective is to name the technology. "Apply machine learning to improve case triage" names a method and measures nothing.

A measurable objective has six parts. Leave one out and acceptance turns into an argument.

Baseline. What the current process produces today, measured the same way the new system will be measured. If nobody has that number, producing it is task one.

Metric. A named quantity with a written formula. Median days to first decision. Share of extracted fields matching the source document exactly. False-negative rate at a fixed alert volume.

Threshold. Two numbers: the value that counts as success and the value below which the deliverable is rejected. A single target invites a long conversation about "close enough."

Instrument. The data the metric is computed on. Name the sample, its size, its date range, and how it was drawn. A metric with no named sample is one the vendor gets to choose.

Scorer. Who computes the number and who checks it. Where judgment is involved, name two reviewers and say how disagreements are settled.

Date. When the measurement happens, and how long the buyer has to review it.

Put together, it reads like this. "On a stratified random sample of 500 closed cases from FY2024 and FY2025, drawn by the program office and withheld from the contractor, the delivered system shall extract the eight fields in Attachment B at a field-level exact-match rate of at least 92 percent. Below 85 percent, the deliverable is rejected. Two government reviewers score it, with disagreements adjudicated by the COR. Measurement occurs at the end of Task 3, with 10 business days for the government to accept or reject in writing."

That paragraph is boring, and it is the difference between a project that ends and a project that renews forever.

That paragraph is boring, and it is the difference between a project that ends and a project that renews forever.

The sections a working scope needs

Order matters little. Presence does. Sections four, seven, eight, twelve, and thirteen are missing from most first drafts, and they are the expensive omissions.

  • Background and current state
  • Objective, written with the six parts above
  • Scope boundaries, including a short list of what is explicitly out
  • Furnished data, systems, and accounts, each with a delivery date
  • Numbered tasks, each with a named output
  • Deliverables and acceptance criteria, keyed to task numbers
  • Security, data handling, and where processing may occur
  • Intellectual property and data rights, including limits on your data
  • Key personnel, named, with substitution rules
  • Schedule, milestones, and payment triggers
  • Reporting cadence and meetings
  • Change control and the single authorized approver
  • Transition and exit package
  • Assumptions and constraints, each written as a testable statement

Data access is the schedule

The most common way an AI project slips is that the data arrives late. Access needs a security review, an owner's approval, a provisioned account, and an open network path. None of that is technical work and all of it takes weeks.

The clause that saves the most money

Treat data access as an obligation with a date

List each dataset, system, and account in the furnished-property section with a delivery date, then state the consequence: if access is late, the period of performance extends day for day and the contractor works the tasks that do not depend on it. That one provision removes the most common argument in the project.

Two more provisions belong here. Name the data owner as a person and a role, not an office, because an office cannot approve anything. And define a fallback that lets work start: a de-identified extract, a synthetic sample on the same schema, or a few hundred records pulled by hand.

If the records fall under the Privacy Act (5 U.S.C. 552a) and the contractor will operate a system of records on your behalf, FAR subpart 24.1 requires the Privacy Act clauses in the contract. Make that call before award; made after, it becomes a modification, and modifications are where prices go up.

Say the security level out loud, before award

Security is where AI pilots die quietly. The work finishes, the demo is good, and nothing can be connected because nobody decided where the data was allowed to live.

Four decisions belong in the scope. What is the sensitivity of the data: Federal Contract Information, CUI, Privacy Act records, health records under HIPAA, student records under FERPA? Where may it be processed: contractor cloud, government cloud, furnished laptops, an isolated enclave? Which control set applies? And may any part of it, including prompts and model outputs, leave that boundary?

For Federal Contract Information, FAR 52.204-21 sets fifteen basic safeguarding requirements. For CUI on a Department of Defense contract, DFARS 252.204-7012 requires NIST SP 800-171 implementation, media preservation, and reporting of cyber incidents to DIBNet within 72 hours. CMMC adds an assessment: the program rule at 32 CFR Part 170 took effect December 16, 2024, and the acquisition rule putting CMMC into DoD solicitations took effect November 10, 2025 on a phased schedule. If the vendor will host anything as a service, FedRAMP authorization is a long-lead item measured in quarters.

The buyer version is short. Pick the level, name it in the scope, and require the vendor to draw the security boundary on one page before work starts. A vendor who cannot draw it has told you something.

Data rights: decide who paid for what

Rights follow funding, and the defaults are more generous to the buyer than most buyers realize. Under DFARS 252.227-7014, noncommercial software developed exclusively at government expense comes with unlimited rights. Mixed funding produces government purpose rights, which convert to unlimited after five years. Private funding produces restricted rights. Technical data follows a parallel structure under DFARS 252.227-7013, and civilian agencies use FAR 52.227-14 with Alternate III for restricted software.

The step that decides all of it happens before award. DFARS 252.227-7017 requires an offeror to assert every restriction on use, release, or disclosure with its proposal, and anything not asserted is delivered with unlimited rights. Read that list closely: it tells you which parts of the system you can modify and which tie you to the vendor. Commercial software is a separate track under FAR 12.212, where the government takes the license customarily offered to the public.

Two clauses are worth adding regardless of contract type. State that your data may not be used to train, tune, or evaluate models serving any other customer, and that it is deleted or returned at close-out with written confirmation. Then require the vendor to name any component whose license would block you from retraining later. OMB Memorandum M-25-22, issued April 3, 2025, directs federal agencies to address both points, and other buyers can borrow the language.

Price discovery separately from the build

This is the structural move that prevents the most rework. Buy the answer to "is this feasible and how would we measure it" as its own small contract, then price the build against what discovery found.

Discovery-then-build, single workflow

1
Access proven and a real extract in hand
1–2 weeks
2
Data profile: completeness, label quality, date coverage
1 week
3
Feasibility test on real records, baseline measured
2–3 weeks
4
Measurable objective, security boundary, build scope with a fixed price
1 week
5
Build against milestone acceptance
8–16 weeks
6
Transition, exit package, reproduction of the accepted metric
2 weeks

Discovery has five deliverables and no others: a data profile, a feasibility result computed on real records, a measured baseline with the proposed objective, an architecture sketch with the security boundary drawn, and a build scope with a fixed price. If a vendor proposes discovery and the deliverable is a strategy deck, that is a different product.

Size it against thresholds you already have. The micro-purchase threshold is $10,000 and the simplified acquisition threshold is $250,000 (FAR 2.101), so most single-workflow discovery is a small purchase and a narrow feasibility test can fit under the micro-purchase line. Commercial services can use FAR subpart 13.5 procedures up to $7.5 million, certified cost or pricing data is not required below $2 million (FAR 15.403-4), and an 8(a) sole-source services award can run to $4.5 million.

One caution first-time buyers miss. Under FAR 9.505-2, a contractor that writes specifications or work statements for a follow-on competitive acquisition generally may not compete to perform that work. Either write the build scope yourself from the discovery output, or put discovery and build in one contract with the build as a priced option exercised only if discovery clears a stated gate. The priced option is usually right: continuity, plus a clean off-ramp.

On contract type: firm-fixed-price with milestone payments fits the build once the objective is measurable, which is the practical argument for discovery first. Time-and-materials is the least preferred type in the FAR (FAR 16.601), so if you use it for discovery, cap the ceiling and attach a deliverable.

Change control: one person can say yes

In federal contracting only the contracting officer can direct a change (FAR 43.102). A program manager's enthusiastic email is not a change, and work performed on it may never be paid for. The rule protects both parties, and it breaks constantly on AI projects because the work is exploratory and the conversations are technical.

Three provisions handle it. Name the contracting officer's representative and state that technical direction within scope comes from that person, while anything touching price, schedule, or deliverables comes only from the contracting officer. Require a written change memo of two pages or less: what changed, why, effect on price, schedule, and acceptance criteria, and the approval line. Keep a numbered change log as a deliverable.

Constructive change claims, where work was directed informally and billed later, are the expensive version of skipping this. The cure is twenty minutes of paperwork per change. State and commercial buyers have no FAR to fall back on, which makes naming the approver matter more.

A reporting cadence that gives you a place to stop

Reporting exists to tell you whether to cancel a project, and when. Four items do it.

  • Weekly one-page status – what moved, what is blocked, what is needed from you, and the current metric value if one exists yet.
  • Biweekly working demo – run live on real or fallback data, no slides. A demo that is a slide deck is a status report in costume.
  • Monthly financial status – funds invoiced and remaining, share of the period elapsed, share of tasks accepted. The gap between the last two is your early warning.
  • Milestone acceptance notice – written acceptance or rejection inside a stated window, any rejection tied to a specific criterion.

Define the review window and say what silence means. Ten business days is normal. Whether no response counts as acceptance is your call; the mistake is leaving it unwritten.

Acceptance and the exit package

Acceptance criteria should copy the objective rather than restate it. If the objective carries the six parts, acceptance writes itself.

The exit package is what buyers forget, and it separates a system you own from a permanent dependency. Require all of it: source code and build scripts in a repository you control; an environment definition that runs on a clean machine; model artifacts with the configuration that produced the accepted result; the evaluation script and the held-out sample; a lineage note stating what data was used for training, tuning, and testing; a runbook covering deployment, retraining triggers, and failure modes; and a software bill of materials, which federal buyers have asked for since Executive Order 14028 in 2021.

Then add the step that makes the scope enforceable. Require your own engineer or a third party to reproduce the accepted metric from that package as a condition of acceptance. Reproduction is the only claim check that does not rest on the vendor's report of the vendor's work.

Common scope failures and what they cost

FailureHow it reads in the scopeWhat it costs
Objective names a technology"Apply machine learning to improve triage"Final review has no test to run. You pay for the dispute, then for a second contract to define what was bought.
No baseline"Improve accuracy of the current process"Improvement cannot be shown. The result is undefendable in a budget review even when the work was good.
Vendor picks the test data"Contractor shall evaluate system performance"Numbers are selected rather than earned. Re-evaluation on a real sample, often a rebuild.
Data access with no date or owner"Government will provide access to relevant data"Weeks of idle burn. On time-and-materials you pay for the waiting.
Security level settled laterSilentA finished system that cannot be connected. Rehost, re-document, or shelve. Usually shelve.
Rights left to the defaultSilent, or a vendor license attached at the backRetraining and modification require the original vendor. Sole-source dependency for the life of the system.

Bottom line

The scope is the first product you are buying. Write the objective with six parts, attach dates to data access, settle the security level before award, read the rights assertions, name one person who can approve a change, and require reproduction of the result from an exit package you hold. Those six keep the negotiating from happening after the money is gone.

Frequently asked questions

How long should a scope of work for AI consulting be?

Six to twelve pages for a single workflow. Section coverage is the measure, not length: a four-page scope missing data access, security, rights, and change control costs more than a twelve-page scope that covers them.

Should I pay for discovery or ask vendors to do it for free?

Free discovery is sales work, and the output belongs to the seller. Pay for it, own the deliverables, and keep the right to use them in a later competition. Most single-workflow discovery fits well under the $250,000 simplified acquisition threshold.

What if a vendor will not commit to an accuracy target before seeing the data?

That is the honest answer, and it is why discovery is priced separately. The target gets committed at the end of discovery, once both sides have seen the real records.

Who owns the model at the end of the contract?

It follows who funded the development. Under DFARS 252.227-7014, government-funded software carries unlimited rights, mixed funding carries government purpose rights that convert to unlimited after five years, and privately funded software carries restricted rights. Restrictions must be asserted with the offer under DFARS 252.227-7017, so read that list before award.

How do I verify that a reported accuracy number is real?

Require reproduction. Hold back a sample the contractor never sees, require the evaluation script and model artifacts in the exit package, and make an independent re-run of the metric a condition of acceptance.

1 business day response

Writing a scope for AI work?

We write measurable objectives, discovery scopes, and acceptance criteria for federal, state, and commercial buyers, and we build the systems those scopes describe.

CapabilitiesMore insights →Start a conversation
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE