Every analytics company that considers federal revenue asks the same first question, and almost nobody answers it properly. The question is how big the market actually is for what this company sells. The bad answers come from analyst reports that size "government IT" at a number so large it is useless for planning. The good answer is a number the company computes itself, from a public record of every federal contract award, filtered to the categories it can actually serve. That record is free, queryable, and updated continuously, and a competent analyst can produce a defensible market size from it in about a day.
This is written for the finance or strategy leader who has to defend a federal investment to a board. What follows is the sizing method, the categories that matter, some anchor figures we pulled from the public record while writing this, and the specific places the method misleads you if you are careless.
The source of record
Federal spending is reported through USAspending.gov, which publishes prime award and subaward data drawn from agency financial and procurement systems. It has a public API and a bulk download facility. Anyone can query it. There is no license, no vendor, and no seat cost, which means the number you produce is auditable by anyone at your board table who wants to check it.
Two organizing codes do most of the work. The North American Industry Classification System code describes the industry of the work being bought. The Product and Service Code describes the thing being bought. Contracting officers assign both at award. Neither is perfectly applied, and the gap between them is where careless sizing goes wrong, which is the subject of a later section. Used together, they bracket the truth.
Anchor figures, and where they came from
We ran the queries below against the USAspending.gov API while writing this article, covering fiscal year 2024, which runs October 1, 2023 through September 30, 2024, and restricted to contract awards. These are the totals the system returned. Anyone can reproduce them.
By industry code, according to USAspending.gov data for fiscal year 2024, computer systems design services, NAICS 541512, accounted for about $33.4 billion in federal contract obligations. Other computer related services, NAICS 541519, accounted for about $25.3 billion. Custom computer programming services, NAICS 541511, accounted for about $10.8 billion. Administrative and general management consulting, NAICS 541611, accounted for about $15.0 billion.
By product and service code, the same USAspending.gov query for fiscal year 2024 shows that the business application and application development support services labor category, PSC DA01, accounted for about $20.8 billion, while the corresponding software category, PSC DA10, accounted for about $6.3 billion. Information technology management support services labor, PSC DF01, accounted for about $4.0 billion.
And by buyer, the same source shows that within NAICS 541512 alone in fiscal year 2024, the Department of Defense obligated about $8.5 billion, the General Services Administration about $6.2 billion, the Department of Veterans Affairs about $4.6 billion, the Department of Health and Human Services about $3.5 billion, and the Department of Homeland Security about $2.6 billion.
Those are anchors, not your market. Your market is a subset of them, and computing that subset accurately is the actual exercise.
What makes a federal sizing exercise trustworthy to a board
Editorial weighting, illustrative rather than measured. The last row is low because a forecast nobody can reproduce is the least useful input in the file.
The sizing method, step by step
The method has six steps. It takes an analyst who can write a Python script about a day, and a team without one about three days using the site's own search interface and downloads. Either way, write down each filter you apply and why, because the defensibility of the answer lives in that list.
Step one: define the thing you sell in the government's vocabulary
This is the step people skip and it determines everything downstream. Write one sentence describing what an agency would be buying from you, using words a contracting officer would use. A recurring data feed with entity resolution. A hosted analytics application. A risk scoring service. Advisory analytics delivered as reports. Each of those maps to different codes, and a company that sells two of them has two markets, not one.
Step two: pull the industry and product codes that contain that thing
Search the public code lists for terms that describe your product, and collect every code that could plausibly carry it. Be generous here. You will narrow later, and a code you never considered cannot be found in step five. The information technology and telecommunications product codes and the professional services codes carry most analytics work, but subscription and reference data sometimes sit in unexpected places, and geospatial products have their own homes.
Step three: query the record for each code, by year and by agency
Pull three years, not one. A single year mixes ordinary spending with the tail of one very large award, and a three-year series shows you whether a category is growing, flat or an artifact. Break each code out by awarding agency, because your entry strategy is a choice of two or three agencies rather than a choice of a market.
Step four: read fifty actual awards
This is the step that separates a real answer from a spreadsheet. Pull the award descriptions for the largest fifty awards inside your codes and read them. You will find that a meaningful share of what a code contains is not what you sell, and that a category you dismissed contains exactly what you sell under a different name. Adjust the codes and rerun. Most companies revise their code list substantially after this step, and the revision is the most valuable output of the whole exercise.
Step five: filter to serviceable
Now subtract. Remove awards to categories of business you are not, awards under a dollar threshold you would not pursue, awards on vehicles you cannot access without a multi-year effort, and awards for work adjacent to yours that you would have to build a new capability to perform. What remains is the serviceable market, and it is usually one to five percent of the headline category number. That collapse is not bad news. It is the difference between a plan and a slide.
Step six: attach the calendar
Federal money is not available continuously. It becomes available when a contract ends and is recompeted, and the end dates are in the public record. Pull the period of performance end date on every award in your serviceable set and build a calendar. That calendar, not the market size, is what determines when revenue can start, and it is the single most useful artifact this exercise produces.
Where the method misleads you
Six traps, each of which we have watched a competent team walk into.
| Trap | What it looks like | The correction |
|---|---|---|
| Obligations read as revenue | A multi-year award appears entirely in the year it was signed | Use outlays or per-year obligations for a run-rate; use total obligations only for deal size |
| One award distorts a category | A code doubles in a year and halves the next | Pull three years and inspect the largest awards before trusting any trend |
| Codes are applied loosely | Your product appears under four codes and dominates none | Bracket with both industry and product codes and reconcile by reading descriptions |
| Vehicles are invisible in the totals | The money is real but only buyable through a contract you do not hold | Tag each award with its vehicle and subtract what you cannot reach this year |
| Subawards are missed | The prime is a systems integrator and your work is a line inside it | Query the subaward data as well; it changes who your customer is |
| Grants are ignored | Analytics bought by a state using federal grant money is invisible in contract data | Size the grant-funded channel separately; it is a different sales motion |
The first trap is the expensive one. Contract obligations record money legally committed, which for a five-year award can all land in one fiscal year in the data even though the work is spread across five. A company that reads obligations as annual revenue overstates the market by a factor that varies by category and can be large. For run-rate purposes, use outlays, which record money actually paid, or compute obligations per year of period of performance.
What a sizing exercise should hand back to a board
Editorial weighting, illustrative rather than measured. The last row is low because a category total is the least decision-relevant number the exercise produces.
What the record tells you beyond the size
Three things, each more useful than the total.
Who the incumbents are. Every award names its recipient. For any category you care about, you can produce a ranked list of the companies currently holding the work, with their contract values and end dates. That is your competitive set, and it is knowable exactly rather than approximately. It also tells you which of them might be a route to market rather than a rival, because a systems integrator holding a large award needs data and analytics components inside it.
How the buying is structured. Award records show the vehicle and the competition status. A category bought overwhelmingly through one government-wide acquisition contract is a category with an access requirement attached, and knowing that before you plan is worth more than knowing the size. A category with many small direct awards is a category you can enter without a vehicle at all.
Where the same buyer buys repeatedly. Group awards by the buying office rather than the agency. Federal purchasing is done by offices with their own habits, and an office that has bought analytics five times in three years is a warmer target than a department with a larger total. The office code is in the data.
An illustrative worked example
Take an analytics company that sells a hosted risk scoring application to commercial banks and wants to know whether federal is worth a two-year investment. Illustratively, the exercise runs like this.
The company's product maps to two product codes and three industry codes. Those codes together carry a very large headline number, most of which is application development labor rather than software. Reading the largest fifty awards, the team finds that roughly a fifth of what the codes contain resembles what they sell, and that a second code family covering subscription information services contains a set of awards that look exactly like their product sold under a different name. They add it.
Filtering to awards above a floor they would pursue, in the four agencies whose missions match the product, and removing the awards riding a vehicle they cannot reach within two years, leaves a serviceable set. The calendar on that set shows a cluster of end dates in a single quarter eighteen months out, and two awards ending sooner. The plan writes itself from there: the two near-term ones are the entry attempt, the cluster is the real prize, and the eighteen months is exactly the time needed to complete the security authorization work that would otherwise disqualify them.
Notice what the exercise produced. Not a market size. A target list, a calendar, and a deadline for the engineering work. That is what a board can approve.
The cost side, which nobody sizes
A federal revenue plan that estimates only the revenue is half a plan. The cost of entry for an analytics company is mostly engineering, and it is estimable in the same week the market is sized.
Deployment engineering. If agencies will require the product to run inside their own boundary, the work is separating the control plane from the data plane, removing or making optional every outbound dependency, replacing managed services with no government-region equivalent, and producing a deployment artifact a government engineer can install without vendor credentials. Size it by inventorying the product's outbound network calls and external dependencies. That inventory takes a week and produces a defensible estimate.
Security authorization. A system security plan against the NIST SP 800-53 control catalog, evidence for each control, a plan of action for gaps, and a continuous monitoring commitment. The engineering half is identity federation, encryption with validated modules, logging and vulnerability management. The writing half is substantial and only engineers can do it accurately.
Accessibility. Section 508 makes accessibility part of what the government may buy, and an accessibility conformance report is frequently required in the proposal itself. Remediating a mature interface costs several times what building it in costs. Start it first because it is independent of everything else.
Data rights review. If your product includes licensed third-party data, someone has to confirm that your licenses permit government use, government hosting, and the retention the government will require. This is a legal review that can invalidate a market, and it is cheap to run early and expensive to discover late.
The federal configuration as a maintained thing. One engineer accountable for it, and a continuous integration job that keeps it working. Without that, the government build breaks two releases after launch and nobody notices until an agency does.
How we work with a company running this decision
Precision Federal builds AI, data platforms, software and cloud systems and delivers them into production, including inside federal agencies. When a commercial analytics company is deciding whether to enter, we do two things and they usually run together.
We run the sizing with your team and hand you the queries. Not a report. A repository: the query scripts, the code lists with the reasoning for each inclusion, the three-year series, the award-level detail for your serviceable set, the incumbent list with end dates, and the recompete calendar. Your analysts rerun it every quarter without us. Companies that use it to decide not to enter have gotten full value from it, and we say so at the start.
We estimate and then build the entry engineering. The dependency and egress inventory, the deployment work, the identity and logging work, the accessibility remediation, the software bill of materials in the build, and the security documentation written by the engineers who built the system. Delivered against acceptance criteria written as tests: an installation that runs from a clean checkout into a government-region environment, an interface that passes an accessibility audit, a control implementation summary an assessor can work from.
What you keep. The code, assigned in writing and committed to your repositories from day one. The data, which we never hold. The customer relationships, entirely. The sizing repository and the queries inside it. Our pre-existing tooling is named in the agreement, excluded from the assignment, and licensed to you perpetually so nothing we bring can strand a future maintainer.
How it is priced. Fixed-price milestones where scope is knowable, which covers the sizing work and most of the deployment and documentation package. Or a committed team for a stated number of months where the work is exploratory. Most companies start with the first.
The first step is one email with a one-page brief: what you sell, who buys it commercially, which agencies you think are candidates, and the date a decision has to be made. We return a scoped, priced statement of work.
Bottom line
The federal market for analytics is large enough that no company needs a forecast to justify looking, and the public record is complete enough that no company should accept a number it cannot reproduce. Pull the categories, read fifty real awards, filter to what you can actually deliver, and attach the contract end dates. The output is not a market size. It is a target list with a calendar, and next to it an engineering estimate for the work that has to be finished before those dates arrive. A board can approve that. A board cannot approve a category total.
Frequently asked questions
Query USAspending.gov, which publishes every federal contract award with an industry code, a product and service code, the awarding agency, the recipient and the period of performance. It has a public API and bulk downloads at no cost. Pull the codes that contain your product for three fiscal years, break them out by agency, then read the largest awards to check that the codes contain what you think they contain. The result is reproducible by anyone.
Obligations are money legally committed at award, which for a multi-year contract can appear entirely in the year it was signed. Outlays are money actually paid out in a period. Reading obligations as annual revenue overstates a market, sometimes badly, because one large multi-year award inflates a single year. For run-rate estimates use outlays or divide obligations across the period of performance. Use total obligations only when you want deal size rather than annual spend.
Far less than the headline. After removing work you do not perform, awards below a size you would pursue, awards restricted to categories of business you are not, and awards riding vehicles you cannot access, the serviceable set is commonly a small single-digit percentage of the category total. That is normal and it is the point of the exercise. The useful output is the named award list and its contract end dates, not the percentage.
Yes. Every award names its recipient, its value, its awarding office and its period of performance end date, and subaward data shows work flowing below a prime. For any category you can produce a ranked incumbent list with expiry dates. That list is your competitive set and also your partnership set, because integrators holding large awards need data and analytics components inside them and buy those components from companies like yours.
Mostly engineering, and it is estimable. Deployment work so the product can run inside an agency boundary without reaching the vendor, identity federation and validated encryption, immutable logging, accessibility remediation in the component library, a software bill of materials generated by the build, and a security documentation package written against the government control catalog. Size it by inventorying every outbound dependency and network call in the product; that inventory takes about a week.
