The interval answers a question nobody in the room asked
The complaint is common enough to be a genre. A team does careful work, reports 0.83 with a 95 percent interval of 0.79 to 0.87, and watches the slide that reaches the decision meeting say 0.83. The interval was not rejected. It was never engaged with, because it describes an estimation procedure applied to a parameter while the reader was choosing between two actions with different costs. No amount of care on the first question answers the second.
Three mismatches do most of the damage, and the engineer can fix all three alone. The first is the quantity: a confidence interval on a mean says where a population average plausibly sits, while the decision turns on the next case, the next batch, the next shipment. The second is units. The interval arrives in F1 score or milliseconds or milligrams per liter; the decision is denominated in reviewer-hours, dollars, schedule risk, or whether a system runs unsupervised. The third is that no action is attached, and an interval that never touches a threshold gets summarized down to its midpoint by anyone under time pressure.
What follows is the translation layer. None of it is new mathematics. Most was standardized decades ago by metrologists, forecasters and intelligence analysts, who faced the same problem under sharper consequences: how to hand a number with error bars to somebody who must act on it and does not share your training.
What a 95 percent confidence interval actually claims
The frequentist definition is a property of a procedure, not of the interval in front of you. Repeat the experiment many times, build an interval the same way each time, and 95 percent of them contain the true value. That says nothing directly about this one, and the gap matters because the intuitive reading is the one everyone uses.
Hoekstra, Morey, Rouder and Wagenmakers put numbers on how deep that runs. In work published in Psychonomic Bulletin & Review in 2014, they gave 120 researchers and 442 students a confidence interval and six statements about it. All six were false. Both groups endorsed, on average, more than three of them. When the people with formal training misread the object at that rate, an interval delivered without translation is not a communication. It is a compliance artifact.
The American Statistical Association reached the same conclusion from another direction. Its statement on statistical significance and p-values, released March 7, 2016, set out six principles, one of which applies directly: "Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold." It also named the alternatives statisticians reach for, including "confidence, credibility, or prediction intervals" and "decision-theoretic modeling." The profession's own guidance treats the interval as a step toward a decision framework, not a substitute for one.
Five intervals that look identical on a slide
Half of the ignored-interval problem is that the engineer computed the interval that was easiest rather than the one matching the question. They render the same way on a chart, so nobody catches it.
| Interval | What it covers | The question it answers |
|---|---|---|
| Confidence interval | A population parameter: a mean, a rate, a difference between two arms | Where does the average sit? Good for comparing systems, weak for the next case, and the one reported when nobody asked which was wanted. |
| Prediction interval | A single future observation, including the irreducible spread across cases | What range should the next unit fall in? What most operational decisions need, and always wider than the confidence interval on the same data. |
| Credible interval | The parameter, under a stated prior and the observed data | Given what we assumed and saw, where do we believe the value is? It supports the probability statement people already think they are reading, at the cost of writing the prior down. |
| Tolerance interval | A stated proportion of the whole population, at a stated confidence | Where do 99 percent of cases fall? The natural form for specification limits, acceptance sampling and performance requirements. |
| Conformal prediction set | The true label or value for a new case, at a chosen error rate | What is the smallest set of outcomes containing the truth 90 percent of the time? Model-agnostic, and the closest thing to a guarantee a learned system offers. |
Picking from this table before the analysis runs is a five-minute conversation that changes what the report can support, because shipping the narrow interval when the decision needs the wide one is an error no reviewer catches from the summary statistics.

Vagueness is what costs credibility, not numbers
Engineers soften uncertainty statements because they believe error bars weaken the result. There is direct evidence on that, and it points the other way.
Van der Bles, van der Linden, Freeman and Spiegelhalter ran five experiments with 5,780 participants, including a preregistered replication with a national sample and a field experiment on the BBC News website, published in PNAS in 2020. They compared numeric uncertainty, stated as a range, against verbal uncertainty and against a control. Numeric ranges produced no significant decrease in trust in the numbers or in the trustworthiness of the source. Verbal uncertainty reduced both.
That should change how a technical report is written. "Approximately 12 percent, though there is some uncertainty in these figures" is the expensive sentence. "12 percent, with a range of 10 to 14 percent" is not. A stated range is close to free. The vague hedge is what costs you the room.
Metrology standardized this a long time ago
Measurement science solved the handoff problem long before anyone had a machine learning model to argue about, because calibration certificates cross organizational boundaries and must mean the same thing on both sides. NIST Technical Note 1297, by Taylor and Kuyatt, splits uncertainty by how it was evaluated rather than by what causes it. A Type A evaluation comes from statistical analysis of a series of observations. A Type B evaluation rests on scientific judgment using all relevant information available: instrument specifications, prior data, published values, knowledge of the process. Both are real, both are expressed as standard deviations, and both enter the combined standard uncertainty.
That split is the one machine learning reporting skips. Variance on a held-out test set is Type A, and it is the only piece most reports contain. That the deployment population differs from the evaluation population, that labels were adjudicated by two annotators rather than five, that the sensor gets recalibrated in the field: those are Type B, and they do not become zero for being uncomputable from the test set. Leaving them out does not produce a conservative estimate. It produces a confident one.
The reporting convention is the second half. Expanded uncertainty is U = k uc(y), where uc is the combined standard uncertainty and k a coverage factor. NIST uses k = 2 by convention, giving a level of confidence of approximately 95 percent for an approximately normal distribution. Because the factor is stated, any reader can divide it back out and recombine the uncertainty with their own. That is what makes a number portable across an organizational boundary: give the value, the interval, the factor, and the method.
The deliverable is a decision rule, not an interval
JCGM 106:2012, on the role of measurement uncertainty in conformity assessment, addresses this most directly. It exists because "9.98 against a limit of 10.00" is not by itself an accept-or-reject answer, and pretending otherwise pushes the risk onto whoever reads the certificate.
Its vocabulary is worth borrowing exactly. A decision rule documents how measurement uncertainty will be accounted for in accepting or rejecting an item against a specified requirement. An acceptance limit is the bound on permissible measured values, which need not equal the tolerance limit. A guard band is the interval between a tolerance limit and the corresponding acceptance limit. The simplest rule, simple acceptance or shared risk, sets the two limits equal and splits the consequences of error between producer and consumer.
Then comes the sentence that belongs on the wall of every acceptance-test review. Under simple acceptance, with a symmetric unimodal distribution for the measurand, the probability of accepting a nonconforming item or rejecting a conforming one can be as large as 50 percent when the measured value lies very close to the tolerance limit. About half the distribution sits on either side of the line, so the call is a coin flip either way. A report reading "0.900 measured against a 0.900 threshold, pass" has hidden a fifty-fifty outcome behind a green cell.
Guard banding is the fix, and it is a business decision wearing statistical clothes. Offsetting the acceptance limit inside the tolerance limit cuts the chance of accepting something nonconforming and raises the chance of rejecting something acceptable. Which error you buy down depends on what each costs your customer, not on the data. The engineer computes the trade curve and names it. The decision-maker picks a point on it.
JCGM 106 also separates two risks that get conflated constantly. Specific risk is the probability that a particular accepted item is nonconforming. Global risk is the probability that a nonconforming item will be accepted on a future measurement. Program offices mean the first when they ask how sure we are about this one. Dashboards report the second. Answering the wrong one sounds like an answer and is not.
Threshold probability is a number the decision-maker already owns
The fastest route out of the trap is to stop reporting where the estimate sits and start reporting the probability that it crosses the line the decision turns on. For that you need the line, and the decision-maker usually knows it even if it has never been written down.
The cost-loss model gives the cleanest version. If a protective action costs C and the loss it averts is L, acting is worthwhile when the probability of the adverse event exceeds C / L. A cheap action against an expensive failure justifies acting on a low probability; an expensive action against a mild failure demands near-certainty. Asking "at what probability would you act?" converts your interval into their units in one question.
Medicine formalized the same idea. Vickers and Elkin introduced decision curve analysis in Medical Decision Making in 2006, defining threshold probability as the minimum probability of an event at which a decision-maker would take a given action, and computing net benefit as a weighted combination of true and false positives with the weight derived from that threshold. Report performance across the plausible range of thresholds rather than at whatever operating point flattered the model, and a buyer can find their own threshold on the curve.
The public example everyone already trusts works the same way. The National Weather Service probability of precipitation is the probability that a given forecast point receives at least 0.01 inch of rain. It is not the fraction of the period during which rain falls, nor the fraction of the area getting wet. That precision is why a 40 percent value is actionable, and it is a fair standard for your own probabilities: state the event, the location, the window and the threshold.
How far each form of uncertainty statement carries a decision
Editorial ranking of how much of a decision each form carries alone. An ordering of usefulness, not a measured statistic.
The ordering is a judgement about what a reader can act on without further work. The top two rows require the threshold to have been obtained from the customer, which is the step that gets skipped. The fifth ranks low not because a confidence interval is wrong but because it leaves the translation to the reader, who resolves it by taking the midpoint. The last ranks lowest because it is the phrasing the trust research found costly.
Calibration is what makes a probability worth quoting
A probability is calibrated when events assigned 70 percent happen about 70 percent of the time. Without that, threshold rules are built on sand: the customer sets a 60 percent action threshold, the model's 60 percent is really 35 percent, and the rule misfires in one direction forever.
The measurement tools are old and well understood. Brier's scoring rule for probability forecasts appeared in Monthly Weather Review in January 1950 and remains the standard summary score. The reliability diagram, predicted probability on one axis and observed frequency on the other, shows in one glance whether the numbers mean what they say. Both belong in any report quoting probabilities.
Modern models make this harder, and the field knows it. Guo, Pleiss, Sun and Weinberger showed at ICML in 2017 that modern neural networks, unlike those of a decade earlier, are poorly calibrated, and identified depth, width, weight decay and batch normalization as factors driving the drift toward overconfidence. Their remedy is temperature scaling, a single-parameter variant of Platt scaling fit on a validation set, effective across most of their datasets. It is about an afternoon of work and it decides whether a threshold rule can be trusted.
One caution belongs beside every calibration claim. Calibration is a population property, and a model calibrated overall can be badly miscalibrated on the subgroup a decision touches: one facility, one shift, one equipment generation. Report reliability per stratum that matters. Calibration also decays as the input distribution moves, a monitoring problem as much as a modeling one; we have written separately on building drift alerts an on-call engineer will still trust.
Coverage you can promise, and the assumption it rests on
Conformal prediction produces the thing decision-makers keep asking for and models rarely give: a stated error rate. It wraps an existing model, uses a held-out calibration set, and returns prediction sets or intervals with finite-sample coverage of 1 − α under an exchangeability assumption, with no assumptions about the data distribution or the model.
The caveats matter as much as the guarantee, and stating them is what makes it survive review. Coverage is marginal: it holds on average over the joint draw of calibration and test data, not conditionally at a fixed input. Distribution-free conditional coverage is not achievable, and adaptive variants that widen or narrow by covariates improve conditional behavior without promising it. Exchangeability is load-bearing, and a shifted deployment population breaks it, which is the Type B uncertainty from the metrology section under another name.
In a report that reads: "Prediction sets contain the true class 90 percent of the time, averaged over cases drawn from the same population as the calibration set. Coverage is not guaranteed for any individual case and degrades if the case mix moves away from that population; the monitoring plan tracks exactly that." Every clause is usable, and none of it is a hedge.
Give them words they are allowed to repeat
Numbers reach the decision table and then get spoken aloud, where they turn back into words. If you do not fix the vocabulary, someone else will, and "likely" leaves the room meaning 51 percent to one participant and 90 percent to another. Two institutions solved this in public, and both solutions are free to copy.
Intelligence Community Directive 203 requires analytic products to express likelihood on a fixed ladder with numeric ranges: almost no chance 01–05 percent, very unlikely 05–20, unlikely 20–45, roughly even chance 45–55, likely 55–80, very likely 80–95, almost certain 95–99. It also enforces a separation worth adopting anywhere: a product expressing the analyst's confidence in a judgment through a confidence level must not combine that confidence level and a degree of likelihood in the same sentence. Two axes, two sentences. How likely the event is, and how strong the basis for saying so is, are different claims, and merging them produces a statement no reader can decompose.
The IPCC's calibrated language does the same work for a different audience. Likelihood terms carry published numeric ranges — virtually certain 99–100 percent, very likely 90–100, likely 66–100, about as likely as not 33–66, unlikely 0–33, very unlikely 0–10, exceptionally unlikely 0–1 — while confidence in the evidence is reported on a separate qualitative scale reflecting the type, amount, quality and consistency of that evidence and the agreement across it.
For an engineering report the adoption cost is one table in the front matter and the discipline to use nothing outside it. Then the words in the summary and the numbers in the appendix are the same object, and a decision-maker quoting the summary is quoting the analysis rather than an impression of it.
Which uncertainty more data will fix
The last translation ends arguments. Split the uncertainty into the part that shrinks with more data and the part that does not. Epistemic uncertainty comes from not knowing enough: too few samples, an under-determined parameter, an unmeasured covariate. Aleatoric uncertainty is the genuine variability of the process, and no dataset removes it.
The decision-relevant question follows immediately, and it is worth asking out loud in the review: would a narrower interval change the action? If the whole interval already sits on one side of the threshold, more measurement buys delay and nothing else. If it straddles the threshold, name the data that would narrow it, what it costs, how long it takes, and how much narrowing to expect. That turns "we need more data" from a request into a priced option, which is the form a program manager can approve.
What belongs in a report that gets used
Eight items, none requiring new analysis. They require only that the work already done be reported in decision form.
- The decision the number is for, in one sentence at the top, including the action it would trigger.
- The threshold, obtained from the customer rather than assumed, and the probability of crossing it.
- The right interval for the question — prediction, tolerance or conformal when the decision is about individual cases.
- The coverage factor, confidence level or error rate, stated so the reader can recombine it with their own.
- The uncertainty that was judged rather than computed, as its own line rather than omitted.
- A calibration check, per stratum that matters, if the report quotes probabilities.
- The decision rule: acceptance limit, guard band if any, and which error it reduces.
- Whether more data would change the action, with the cost and the expected narrowing.
What a buyer can already require in writing
Buyers do not have to invent any of this. The instruments governing the work already reach it, and a supplier writing in these terms is answering the customer's own documents.
| Instrument | What it says | What it means for uncertainty |
|---|---|---|
| FAR 37.601(b) | Performance-based contracts shall include a performance work statement, measurable performance standards in terms of quality, timeliness and quantity, and the method of assessing contractor performance | The method of assessment is a contract term, which makes the decision rule one too. A threshold with no stated treatment of measurement error is an incomplete standard. |
| OMB M-25-21 (April 3, 2025) | Directs documenting known capabilities and limitations of the AI, ongoing testing and validation including testing in real-world conditions, and assessing for overfitting to known test data | Limitations are a required deliverable, not a disclosure the supplier elects to make. Real-world testing is where Type B uncertainty becomes measurable. |
| OMB M-25-22 (April 3, 2025) | Names performance-based techniques for AI acquisition: statements of objectives and performance work statements, quality assurance surveillance plans, and contract incentives tied to performance | The surveillance plan is the natural home for the acceptance limit, the guard band and the monitoring thresholds, negotiated before award rather than argued after delivery. |
| NIST AI RMF 1.0 (AI 100-1) | The MEASURE function applies quantitative, qualitative or mixed methods to analyze, assess, benchmark and monitor AI risk, with associated measures of uncertainty | Uncertainty is named as part of the measurement. A benchmark table without intervals does not satisfy the function it claims to implement. |
| JCGM 106:2012 | Defines the decision rule, acceptance limits, guard bands, and the separation of specific from global consumer and producer risk | Ready-made vocabulary for acceptance testing, letting the engineering argument and the contracts argument use the same words. |
Common objections
Our customer does not want probabilities. They want a yes or a no.
They should get one. The decision rule produces exactly that: a documented rule turning a measurement and its uncertainty into accept or reject. What changes is that the rule is written down and its error rates stated, so the yes means something specific and the same measurement gives the same answer next quarter. Withholding the probability does not remove the risk; it moves it onto whoever signs.
Adding intervals will make our results look worse than a competitor's point estimates.
Against an inattentive reader, briefly. Against a reader burned once by a demo number that did not survive production, the opposite happens, and those readers increasingly run the evaluation. The trust research is direct: numeric ranges did not reduce trust in the numbers or the source, while verbal hedging did.
We cannot quantify the biggest uncertainty, so we left it out.
Leaving it out is itself a quantification, and the value assigned is zero. Measurement science handles this with Type B evaluation: assign a standard uncertainty using judgment and available information, state the basis, and combine it. A stated assumption a reviewer can argue with beats a silent one.
Frequently asked questions
Because the interval describes an estimate and the reader is choosing an action. It is usually on the wrong quantity, a population mean rather than the next case, denominated in modeling units rather than cost, and attached to no threshold. Converting it into the probability of crossing a threshold the customer names fixes all three.
A confidence interval covers a population parameter such as a mean. A prediction interval covers a single future observation, so it includes case-to-case variability as well as uncertainty in the estimate. On the same data it is always wider, and it is what most operational decisions need.
A guard band is the interval between a tolerance limit and the acceptance limit actually used for the accept-or-reject decision. It is warranted whenever a measurement near the limit carries real consequences, because under simple acceptance a value sitting on the limit is close to a coin flip. The offset trades the chance of accepting something nonconforming against the chance of rejecting something acceptable, so its size is a business choice about which error costs more.
Check calibration on held-out data: bin the predicted probabilities, compare each bin against observed frequency, and plot the reliability diagram alongside a summary score such as the Brier score. Do it per subgroup that matters. Modern neural networks tend toward overconfidence, and temperature scaling on a validation set is an inexpensive correction.
