Skip to main content
Technology Decisions

A second opinion on a build-versus-buy decision

Most teams asking for an outside read have already picked a side. A second opinion earns its fee by attacking the assumptions underneath the pick, not by restating the two options. Five tests do most of the work.

The decision is usually made before the review starts

By the time an outside read gets requested, the meeting invite still says "build versus buy" but the deck says something else. Someone has a favorite. The favorite has momentum. The analysis that follows gets built to confirm it, and every number in the spreadsheet is chosen by a person who already knows which way the answer should come out. That is not dishonesty. It is how any group behaves under a deadline with a budget cycle closing. It does mean that a second opinion which only lines up the same two columns is worth nothing, because it will be read through the same lens that produced them.

A useful second opinion does something narrower and harder. It finds the assumptions that are load-bearing, ranks them by how much damage they do if they are wrong, and tests the two or three cheapest ones before the commitment is signed. Teams rarely lose money on the choice itself. Both paths can work. They lose money on an assumption nobody wrote down, which held for the pilot and stopped holding in month fourteen.

This is the working method our engineers use when a client brings us a decision that is already leaning. Five tests: differentiation, data ownership, total cost including the maintenance tail, hiring reality, and exit cost. Run honestly, they confirm the leaning about two-thirds of the time. The other third is where the fee pays for itself many times over.

Test one: does this system change what someone pays you for?

Differentiation is the only argument that makes a build defensible on its own. Ask the question in revenue terms, not in engineering terms. If a customer, a source-selection board, or a regulator could not tell the difference between your implementation and a competent commercial one, you are about to spend engineering capital on undifferentiated plumbing. Nobody has ever won a recompete because their identity provider was hand-rolled.

The test has a sharp edge in federal work. Under FAR 15.101-1, a best-value tradeoff source selection lets the government pay more for technical superiority, so a genuinely better component can be worth its cost on the scorecard. Under FAR 15.101-2, lowest-price technically acceptable does the opposite: above the acceptability threshold, extra quality earns exactly zero. Building something excellent into an LPTA bid is a donation. Read the section M evaluation criteria before the build case is finalized, because the same component can be an asset in one procurement and dead weight in the next.

There is also a statutory tilt toward buying. FAR Part 12 and the market-research requirement at FAR Part 10 push agencies and their contractors toward commercial products and services where one will meet the need, and 10 U.S.C. 3453 makes that preference explicit for defense acquisitions. A federal program that builds where a commercial product exists needs a reason on the record. That reason is usually one of three: no commercial product meets the requirement, the data cannot leave the boundary, or the capability is the differentiator being sold.

Where a build case usually holds up

Domain logic specific to your operation
92%
Evaluation suite and acceptance tests
89%
Data contracts and schemas you must control
84%
The screen your users touch every day
76%
Model serving and orchestration
67%
Identity, storage, base infrastructure
58%

Editorial weighting from practitioner reading. Illustrative, not a measured statistic.

Test two: who owns the data, and who owns what the data produces

Data ownership is two separate questions and most contracts only answer the first one clearly. The first is the raw data: your records, your telemetry, your documents. The second is the derived material: trained weights, fine-tuning checkpoints, embeddings, labeled corpora, evaluation sets, tuned configuration. The second category is where the value accumulates, and it is the category vendors quietly keep.

Read the model-improvement clause in any agreement before the build case is compared against it. Three questions settle it. On termination, what do we receive? In what format? Within how many days? A vendor who will return raw data as CSV in thirty days but keeps every label your analysts produced has taken the asset and left you the receipt. That is a legitimate commercial position. It just has to be priced into the comparison instead of discovered in year two.

Federal work adds a rights layer that has to be settled before award, not after. FAR 52.227-14 gives the government unlimited rights in data first produced under a contract unless limited rights are asserted and marked. On the defense side, DFARS 252.227-7013 and 252.227-7014 set the ladder for noncommercial technical data and computer software: unlimited rights when development was funded exclusively by the government, restricted or limited rights when funded exclusively at private expense, and government purpose rights when the funding was mixed. SBIR-developed data carries its own protection under DFARS 252.227-7018, and the SBA SBIR and STTR Policy Directive sets that protection period at twenty years from award. Those rights are real money, and they are lost by paperwork more often than by negotiation.

Before award, not after

The assertions table is part of the decision

If a commercial component ends up inside a federal deliverable, the assertions required by DFARS 252.227-7017 have to list it, with the correct rights category, before award. A component chosen in an architecture meeting and never surfaced to the contracts side is the single most common way a firm gives away rights it intended to keep. Settle the markings while both options are still on the table.

Test three: total cost, including the tail nobody staffs

Most build-versus-buy spreadsheets compare one year of engineering against one year of license. That comparison is wrong on both sides, and it is wrong in opposite directions, which is why it feels balanced.

The buy column usually omits integration. Connecting a commercial product to your identity provider, your data platform, and your existing workflows is real engineering, frequently three to six months of it, and it repeats in smaller form at each major vendor release. It omits price movement: renewal uplifts in the range of five to eight percent are ordinary, seat growth compounds on top, and the pilot rate is almost never the steady-state rate. It omits the professional-services statement of work that appears once the product meets your actual data.

The build column omits the run cost. A working rule from production systems: a system that took N engineer-months to build consumes roughly fifteen to twenty-five percent of that effort every year, permanently, in dependency upgrades, security patching, on-call rotation, data drift, model refresh, and keeping the tests honest. A twelve-engineer-month build is a two-to-three engineer-month annual obligation for as long as the system runs. Very few estimates carry that line, and no estimate we have reviewed carried it past year three.

Model both columns over five years. The GAO Cost Estimating and Assessment Guide (GAO-20-195G) is the standing public reference for life-cycle cost estimating and its structure is worth borrowing even for commercial work, mainly because it forces the estimator to name uncertainty ranges instead of single numbers.

Five-year cost lineUsually missed on the buy sideUsually missed on the build side
Integration and identityThree to six months of engineering before first production useThe same work, plus you own both ends of it forever
Price movementRenewal uplift of 5% to 8% plus seat growth; pilot rate is not the renewal rateNo license line, but cloud and inference spend scale with usage
MaintenanceBundled into subscription and invisible until a version is deprecated15% to 25% of the original build effort per year, indefinitely
Evaluation upkeepYou inherit the vendor's tests, which measure the vendor's prioritiesThe line most estimates leave out entirely
Security and authorizationInherited controls look free until the vendor is replaced and the package reopensYour system security plan, your control implementations, your assessor
ExitExport format, workflow re-implementation, and whatever the termination clause allowsKnowledge loss the week the person who built it takes another job

One accounting fact decides more of these debates than any technical argument, and it deserves to be named out loud. For a commercial company, ASC 350-40 permits capitalizing qualifying internal-use software development, and ASU 2018-15 extended similar treatment to implementation costs in a hosted arrangement. Federal reporting entities follow FASAB SFFAS No. 10 for internal-use software. Where work funded by a federal award is involved, allowability runs through 2 CFR part 200 subpart E. The practical effect is that a capitalized build can look cheaper on the operating statement than a subscription that hits expense every month, even when the cash is identical or worse. If the build case is winning on the P&L presentation rather than on the cash, that is worth knowing before the vote.

Own the tests and you can rent the engine. Rent the tests and you are renting the answer too.

Test four: the hiring plan hidden inside the word "build"

"Build" always means one of two things. Hire, or the people you already have stop doing something else. Both are legitimate. Only one of them usually appears in the plan.

Name the three seats. A machine-learning system in production needs someone who owns data pipelines, someone who owns serving and infrastructure, and someone who owns evaluation. One capable person carries two of those for a while. Nobody carries three for two years without the third one quietly rotting, and evaluation is always the one that rots.

Price the hiring calendar honestly. Opening a requisition for an experienced ML engineer and having that person productive is a multi-month exercise in any market, and the estimate should carry the vacancy months as cost, not as zero. If the role touches classified work, the calendar gets longer: personnel security clearances are measured in months, and the sponsoring company itself has to hold a facility clearance under the NISPOM rule at 32 CFR part 117 before it can sponsor anyone at all.

Check the bus factor before you commit. If the build depends on the one engineer who is already the busiest person in the company, that is not a staffing plan, it is a single point of failure with a calendar invite. The honest version of that plan names a second person who can carry the system, and budgets the time for that person to actually learn it.

Ask what stops. When existing staff absorb the build, something else gets deferred. Write down what. A decision that quietly trades six months of customer-facing roadmap for an internal platform is a real decision and deserves to be made on purpose.

Test five: what it costs to leave

Exit cost is the number that almost never appears in either column, and it is the one that turns a reversible decision into an irreversible one. Compute it in both directions, because both paths have a lock.

Leaving a vendor costs: exporting data in a format that is usable rather than merely delivered, re-implementing the business logic that accumulated inside their configuration screens, retraining every user, and whatever the termination terms permit. For commercial-item contracts, FAR 52.212-4(l) gives the government a termination-for-convenience right; a private buyer has whatever was negotiated, which is often much less. A vendor holding a FedRAMP authorization inside your boundary adds another lock, because replacing them reopens the authorization package and the assessment schedule that goes with it.

Leaving your own build costs differently. The code is yours, but the understanding walks out with the person who wrote it. A build with no design record, no runbook, and no test suite that a new engineer can pass on day one is as locked as any proprietary product, only without a support line to call.

Four protections are cheap to negotiate while both options are live and expensive to obtain later: a tested export in an open format, exercised on a schedule rather than promised in a clause; source-code escrow with a verification provision, since escrow without verification is theater; a cap on annual price increases written into the first term; and, on the build side, documentation as an acceptance criterion instead of an aspiration.

The answer that usually wins

Run all five tests and the recommendation is rarely a clean side. It is a line drawn through the system. Buy the substrate: identity, storage, orchestration, base models, monitoring, everything where a competent commercial product exists and your version would be indistinguishable. Build the thin layer that is actually yours: the domain logic, the data contracts, the interface your users live in, and above all the evaluation suite.

The evaluation suite is the piece worth building almost every time. It is small, it is cheap, and it is the only artifact that lets you change your mind later. With your own test set and your own scoring, swapping a vendor becomes a measurement instead of a leap of faith, and a vendor's next release becomes something you can verify rather than something you absorb. Firms that own their tests negotiate renewals from a different position, because they can say what the product is worth to them in numbers the other side cannot argue with.

Five ways the assumption fails

The demo ran on the vendor's data. Every product looks excellent on the corpus it was tuned against. The only demo that counts is the one run on a slice of your records, with your edge cases, scored by your criteria.

The build estimate has no evaluation line. If nobody costed the test set, the labeling, and the ongoing scoring, the build estimate is short by a meaningful fraction and the system will ship without a way to know whether it works.

The pilot price is not the renewal price. Introductory terms are a customer-acquisition cost on the vendor's side. Model year three at list, with the seat count you expect to have by then.

"We already have the data." Usually the data exists somewhere, in a format nobody has validated, with gaps nobody has counted. Sample a thousand records by hand before anyone builds a plan on top of it.

The champion is the only user. A tool with one enthusiastic advocate and no measured demand from the people who would use it daily is a preference, not a requirement. Count the users before you count the savings.

How we run a second opinion

The engagement is deliberately small and fixed in scope. Our engineers read what already exists, extract the assumptions, test the cheap ones against your data, and return a written read that a board or a program office can act on. Precision Federal is an SBIR and STTR shop led by a former professor in technology who ranks in the Kaggle Top 200 of more than 200,000 competitors and holds seven cloud certifications, with twenty years building production systems for federal agencies across five consulting firms, three of them federal. Behind that sits a standing bench of named engineers, licensed professional engineers, and domain specialists across defense, health, energy, transportation, and public-sector data, so the reviewer on a given decision is someone who has shipped that class of system.

How the review runs

1
Read both options exactly as written, plus the vendor SOW and the build estimate
Day 1
2
Extract every assumption and rank by how much it moves the answer if it is wrong
Days 2-3
3
Test the two cheapest high-impact assumptions on a slice of your real data
Days 4-8
4
Five-year cost model with the maintenance tail, price movement, and exit line
Days 9-10
5
Written read: recommendation, cost to reverse, and what evidence would change it
Day 11

The deliverable names a recommendation. It also names the reversal cost of that recommendation and the specific evidence that would flip it, because a review that cannot be argued with is a review that was not neutral. If the answer is that your leaning was right, the document says so in the first paragraph and spends the rest of its pages on the assumptions to watch.

Bottom line

Build versus buy is not a values question, and it is not settled by which side has the better slide. It is settled by five checks that most teams skip because the calendar is short: whether the system is a differentiator, who owns the derived data, what the five-year cost looks like with the maintenance tail included, whether the hiring plan is real, and what it costs to walk away. A team that runs those five honestly and still lands where it started has bought something valuable: a decision it can defend for three years without flinching.

Frequently asked questions

How is a second opinion different from a vendor evaluation?

A vendor evaluation scores products against requirements. A second opinion tests the requirements themselves, along with the cost model, the staffing plan, and the assumptions holding the preferred option up. The two work well in sequence, with the second opinion first, because it often changes what gets scored.

We are already leaning toward one option. Is a review still worth it?

Yes, and that is the usual case. Roughly two-thirds of these reviews confirm the leaning. The value in those cases is the written list of assumptions to monitor and the reversal cost, both of which matter when the decision gets questioned a year later by someone who was not in the room.

If you also build systems, how is the recommendation neutral?

The review is priced as fixed-fee work, separate from any implementation, and the recommendation ships with the evidence behind it so anyone can check the reasoning. When the honest answer is to buy a commercial product, the document says to buy it and names which one and why.

What does the maintenance tail on a custom build actually cost?

Plan on fifteen to twenty-five percent of the original build effort every year, indefinitely, covering dependency upgrades, security patching, on-call, data drift, and keeping the evaluation suite current. On a twelve engineer-month build that is roughly two to three engineer-months per year that no one has staffed.

Which data rights questions matter most on a federal program?

Who funded the development, what has been asserted and marked, and whether any commercial component sitting inside the deliverable was surfaced to the contracts team before award. FAR 52.227-14, DFARS 252.227-7013 and -7014, and the SBIR data-rights clause at DFARS 252.227-7018 set the ladder; the assertions table at DFARS 252.227-7017 is where rights are usually won or lost.

1 business day response

Send both options and the assumptions behind them

Email [email protected] with three things: the build estimate and its staffing plan, the vendor quote or SOW, and the assumptions each one rests on. One page each is enough. Within one business day you get a straight yes or no on whether an outside read would change anything, and if it is a yes, a one-page scope with a fixed fee and an eleven-day turnaround.

Start a conversationCapabilitiesMore insights →
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE