One service, four numbers, and only one of them is money
A single knee MRI generates at least four dollar figures before anyone has paid anything. There is the billed charge, taken from the hospital's chargemaster and printed on the claim. There is the allowed amount, the contracted rate the plan and the facility agreed to, which is what the claim is actually adjudicated against. There is the paid amount, which is the allowed amount minus whatever the patient owes under deductible, coinsurance and copay. And there is the patient responsibility, the remainder. Three of those four appear on the same 837 claim and 835 remittance pair, in different fields, and a pipeline that grabs the wrong one produces a number that is internally consistent, passes every schema check, and is wrong by 400 percent.
This is the most common silent failure we see in health-cost products. It throws no error and fails no unit test. It shows up months later, when a customer with real claims experience sees your benchmark for a routine outpatient procedure at four times what they pay, and stops trusting the platform. Recovering from that conversation costs more than the original build.
The decision matters because the gap is enormous and it is not a constant. Hospital charge-to-cost ratios in the United States commonly land between 3 and 6 for the average facility, with a long tail of facilities well above 10. Commercial allowed amounts for the same service routinely run 150 to 300 percent of the Medicare rate. Medicare itself pays a fee-schedule amount that ignores the charge entirely. So the relationship between the number on the bill and the number that moved is not a scaling factor you can correct with a multiplier. It varies by facility, by service line, by payer, and by contract year.
Same service, different measures — typical spread against the Medicare rate
Illustrative ranges drawn from published RAND hospital price studies, MedPAC reports and CMS transparency files. Actual spreads vary widely by market, service line and contract.
Where each number actually comes from
The chargemaster is an internal price list maintained by the facility. Its numbers are set by a mix of history, annual percentage increases, and the handful of contracts that still pay a percentage of charges. For most payers the chargemaster is a starting position that gets discarded during adjudication. It survives in the data because it is what the provider submits, and because two categories of people still face it directly: uninsured patients who are billed before any discount policy is applied, and out-of-network situations where no contract governs.
The allowed amount is the contract. It comes from the payer's fee schedule for that provider, and it is the number the plan and the facility agreed represents the price of the service. In claims data it typically appears as an adjudicated line-level figure derived from the submitted charge minus the contractual adjustment. In an 835 remittance that contractual write-off carries claim adjustment reason codes in group CO, and the allowed amount is what remains before patient liability comes out.
The paid amount is what the plan wired. It is the allowed amount minus deductible, coinsurance, copay, and any coordination-of-benefits offset. Paid amount is the right number if the question is what the plan spent. It is the wrong number for almost every price comparison, because two identical services at the same facility on the same contract will show wildly different paid amounts in January and October, purely as a function of where each patient sits against the deductible.
That last point deserves emphasis because it produces a specific, reproducible artifact. Build a benchmark on paid amount, group by month, and you will see a beautiful seasonal curve in your prices. It is not a market signal. It is the deductible cycle, and it will be visible to any customer with an actuary.
What the transparency rules put in reach, and what they did not
Two federal rules changed what is buildable here without buying a claims database. The Hospital Price Transparency rule, at 45 CFR Part 180 and effective January 1, 2021, requires hospitals to publish a machine-readable file of standard charges covering gross charges, discounted cash prices, payer-specific negotiated charges, and the de-identified minimum and maximum negotiated charges. CMS tightened the format substantially, moving hospitals to a required CMS template schema with attestation, and the enforcement posture has hardened from warning letters to civil monetary penalties.
The Transparency in Coverage rule, at 45 CFR 147.210, put the payer side into the open. Group health plans and issuers publish in-network negotiated rate files and out-of-network allowed amount files, monthly, as machine-readable JSON. These are the largest public price files in American healthcare, and they are also the reason many teams underestimate the work. A single national carrier's monthly in-network file set runs into the terabytes. The index file points to hundreds of thousands of blobs. Rates are keyed to provider groups by EIN and NPI, and the same negotiated rate can appear under a dozen billing codes and several contract arrangements, including percentage-of-charge and per-diem entries that are not a dollar price at all.
The important limitation is that neither file set contains volume. A negotiated rate file tells you the price of a service at a facility under a plan. It does not tell you whether that facility performed the service once last year or four thousand times. Any average computed from these files is unweighted, and an unweighted average across a list that includes rates for services a provider never delivers is a number without a referent. Getting to a weighted figure requires joining utilization from somewhere else: a claims database, an all-payer claims database, or a public utilization file such as the CMS Medicare provider utilization and payment data.
All-payer claims databases and the ERISA hole
Roughly twenty states operate an all-payer claims database, and several more have authorizing statutes in various stages. These are the closest thing to a market-wide view of allowed amounts, and where one exists with a workable data release process it is often the best foundation available for a state-level product. Colorado, Massachusetts, Maryland, Minnesota, Virginia, Utah and New Hampshire all run mature programs with published data use agreement processes.
They come with one structural gap that anyone building on them must understand and disclose. In Gobeille v. Liberty Mutual, 577 U.S. 312 (2016), the Supreme Court held that ERISA preempts state reporting requirements as applied to self-funded plans. Since roughly two-thirds of covered workers in employer plans are in self-funded arrangements, that decision removes a large share of the commercial market from mandatory state collection. Congress responded in part through the No Surprises Act, which directed the Department of Labor to establish a voluntary reporting format for self-funded plans, but voluntary is doing real work in that sentence.
The practical consequence for a product is a coverage question you have to answer honestly in your methodology. If your Colorado commercial benchmark is built on an APCD, some self-funded volume is missing, and the missing share is not random. Large employers are more likely to self-fund, and large employers negotiate differently. State off it and your benchmark carries a bias you did not choose and cannot see.
The cost-to-charge ratio trap
When only charges are available, the standard move is to convert them to something cost-like using a cost-to-charge ratio from the Medicare Healthcare Cost Report Information System. HCRIS publishes facility-level ratios, and CMS itself uses departmental ratios in the outpatient and inpatient payment systems. Applying them is defensible, well-documented, and used in serious published research. It is also frequently misunderstood by the teams that adopt it.
A cost-to-charge ratio estimates the facility's accounting cost of providing a service. It does not estimate the price anyone paid. Multiplying a charge by a ratio of 0.25 gives you an approximation of what it cost the hospital, which is a different quantity from the negotiated rate, usually well below it. If your product tells an employer what a procedure should cost and the underlying method is charges times a cost-to-charge ratio, you are showing them hospital accounting cost and labeling it price. That is a defensible research estimate presented as a commercial claim, and the gap between those two things is where product credibility goes to die.
Ratios also degrade. HCRIS reports lag a year or more, are facility-wide unless you use departmental detail, and stop describing a facility that has shifted service mix. If a ratio is in your pipeline, its vintage belongs in your methodology document and its sensitivity belongs in your test suite.
Site of service is the variable people forget
The same CPT code costs different amounts depending on where it happens, and the difference is structural, not noise. A procedure performed in a hospital outpatient department generates a facility fee under the Outpatient Prospective Payment System in addition to the professional fee. The same procedure in a freestanding ambulatory surgical center pays under the ASC schedule, which for many codes sits well below the OPPS rate. In a physician office it pays under the Physician Fee Schedule with a non-facility practice expense component and no separate facility charge at all.
A benchmark that mixes sites of service without controlling for them produces a distribution with two or three modes and a mean that describes nothing. Worse, hospital acquisition of physician practices moves procedures across those categories over time, so a longitudinal series built without site control will show price growth that is partly real and partly a reclassification artifact. The fix is not complicated. Carry place of service code, revenue code, bill type, and the professional-versus-facility split as first-class dimensions, and never aggregate across them without saying so.
| If the question is | Use this measure | Because |
|---|---|---|
| What does this service cost in this market? | Allowed amount, weighted by volume | The contract price is the market price; volume weighting keeps rare-service rates from dominating |
| What will this plan spend next year? | Paid amount plus patient liability, modeled separately | Plan cost and member cost move on different levers; benefit design changes one without the other |
| What will this patient owe? | Allowed amount plus the member's accumulator position | Liability is a function of the contract price and where the deductible stands on the service date |
| How does this hospital price against peers? | Allowed amount as a percent of Medicare | Medicare is the only common denominator across markets with different cost structures |
| What does an uninsured patient face? | Gross charge and posted discounted cash price | These are the only numbers that actually govern a self-pay encounter |
| What did it cost the facility to deliver? | Charges times a departmental cost-to-charge ratio | Accounting cost, explicitly not price; label it that way everywhere it appears |
The join that breaks first: provider identity
Every one of these datasets keys providers differently, and the identity join is where more engineering hours go than anyone budgets. Transparency in Coverage files key to EIN and NPI. Hospital transparency files key to a facility and often a CCN. Claims carry both a billing NPI and a rendering NPI, which are frequently different entities. APCD extracts key to their own internal provider identifiers with a crosswalk of variable quality. NPPES gives you the registry, and it is self-reported, stale in places, and full of organizations that share addresses.
The failure mode is subtle. A health system with 40 NPIs across 12 facilities will match partially, and a partial match produces a benchmark that describes some of the system's locations rather than the system. Nobody notices, because the output looks like a number. Our approach is to build the provider crosswalk as a first-class versioned artifact with its own test suite, resolve to a system-level identity explicitly rather than as a byproduct of a fuzzy match, and carry a match-confidence field all the way to the surface so a low-confidence roll-up can be suppressed instead of quietly displayed.
Expect this to be a genuine work stream. On a national build touching Transparency in Coverage files, hospital MRFs and a claims source, provider identity resolution and validation is commonly 30 to 40 percent of the total engineering effort. Teams that budget it as a two-week task ship late and ship wrong.
Where the engineering hours actually go on a health-cost build
Relative effort weighting from practitioner experience on multi-source price builds — illustrative, not a measured statistic.
Privacy constraints that shape the schema, not just the access policy
Price data is not automatically free of protected health information. A claims-derived benchmark at a granular enough cut becomes re-identifying, and the constraint belongs in the design rather than in a downstream access rule. The HIPAA de-identification standard at 45 CFR 164.514 gives two paths: Safe Harbor, which removes 18 identifier categories and forbids actual knowledge of residual risk, and Expert Determination, which requires a qualified person to document that re-identification risk is very small.
For a product that reports prices by provider and service, Expert Determination is usually the realistic path, and it comes with obligations that reach into the schema. Minimum cell sizes, typically 11 following CMS practice, have to be enforced at query time and not just at load. Complementary suppression is required so that a suppressed cell cannot be recovered by subtraction from a published total. Date granularity gets reduced. Geographic granularity below the three-digit ZIP needs a specific justification. If your APCD data use agreement adds its own release rules, and most do, those stack on top rather than replacing the HIPAA analysis.
Teams that treat suppression as a presentation-layer filter get caught in review. The determination expert will ask whether an API can be queried repeatedly to reconstruct a suppressed cell, and if the answer is yes, the determination does not hold. Build the suppression into the aggregation layer with query-budget awareness, and keep the logic in one place so a new endpoint cannot bypass it.
Write the measure definition before you write the pipeline
The single practice that separates health-cost products that survive scrutiny from those that do not is a written measure specification produced before any code. Not a data dictionary. A specification that states, for each published figure, the numerator, the denominator, the population, the inclusion and exclusion criteria, the unit of analysis, the weighting, the reference period, the risk adjustment if any, and the known limitations.
This is not a documentation nicety. It is the artifact that makes disagreement productive. When a customer says your number is wrong, the specification turns an argument about credibility into a technical conversation about a definition, and definitions can be reconciled. Without it, you are defending a number with no provenance, which you will lose.
The frameworks that govern this kind of work all point the same direction. The NCQA HEDIS measure specification format is the reference standard for how a healthcare measure gets written down. The CMS Measures Management System Blueprint documents the full lifecycle from concept through maintenance. For a model in the pipeline, Federal Reserve SR 11-7 remains the clearest available template for what independent validation looks like: conceptual soundness, outcomes analysis, benchmarking against a challenger, and limitations stated plainly. None of these are binding on a commercial health-cost product. All of them are recognizable to the analytics leaders who evaluate one, and writing in their form is free credibility.
How we scope this work
Our engineers start a health-cost engagement with a measure workshop rather than an architecture diagram. The first session ends with a one-page definition of every published figure, and it usually surfaces a disagreement inside the client's own team about what the product is supposed to be saying. That disagreement is cheap to resolve in week one and expensive to resolve after the pipeline exists.
The second step is a data reality assessment against actual files, not documentation. Pull three real MRFs from carriers you intend to cover, three hospital files from facilities in your target market, and if a claims source is in play, a real extract. Measure the schema variance, the fill rates on the fields your measure depends on, the percentage of rates that are percentage-of-charge or per-diem rather than dollar amounts, and the provider match rate against your crosswalk. That week reprices the project honestly, and it has killed measures never computable at the fidelity the roadmap assumed.
Then the pipeline gets built with the measure definition as its test oracle. Every published figure has a test that recomputes it from a fixed sample by an independent path. Suppression is tested by adversarial query. The provider crosswalk is versioned and diffed on every refresh, with a report of what moved. Typical range for a first production build across two data sources with a defined measure set is 12 to 20 weeks and 900 to 1,600 engineering hours, and the variance is driven almost entirely by how many payers and how many provider systems are in scope, not by analytical complexity.
Bottom line
Charges, allowed amounts and paid amounts answer different questions, and the choice among them is a product decision disguised as a data decision. Allowed amount weighted by volume is the right default for anything describing what a service costs in a market. Paid amount belongs in plan spend analysis and nowhere near a price benchmark. Charges belong in self-pay and out-of-network contexts, and in cost estimation only when the output is labeled as accounting cost. Write the measure definition first, resolve provider identity as a real work stream, build suppression into the aggregation layer, and publish the limitations before a customer finds them. The firms that do this ship products that survive the actuary. The ones that skip it ship a number that nobody can defend.
Frequently asked questions
The allowed amount is the contracted price the plan and provider agreed on for a service. The paid amount is what the plan actually sent, which is the allowed amount minus deductible, coinsurance and copay. Allowed amount is stable across patients on the same contract; paid amount swings with each member's accumulator position.
You can build an unweighted rate benchmark. You cannot build a volume-weighted one, because the files contain no utilization. Without volume weighting, rates for services a provider rarely performs carry the same weight as its highest-volume procedures, which distorts any average.
Because of Gobeille v. Liberty Mutual, 577 U.S. 312 (2016), in which the Supreme Court held that ERISA preempts state reporting mandates as applied to self-funded plans. Reporting by those plans is voluntary, and self-funded arrangements cover roughly two-thirds of workers in employer coverage.
It is a valid way to estimate the facility's accounting cost, which is a different quantity from price and usually well below the negotiated rate. Using it is defensible research practice as long as the output is labeled as cost, never as what a payer or patient pays.
HIPAA de-identification under 45 CFR 164.514 governs, usually through Expert Determination for this kind of product. In practice that means minimum cell sizes around 11, complementary suppression so totals cannot reveal a hidden cell, and enforcement inside the aggregation layer rather than at the presentation layer. Data use agreements from a state APCD add their own rules on top.