Skip to main content
Experimentation & Causal Inference

Causal questions an A/B test cannot answer

Randomization answers one question extremely well. A surprising number of the decisions a company actually makes are a different question wearing the same clothes. Here is how to tell them apart, and what to run when the test is not available.

What a test actually answers

An A/B test answers a narrow question with unusual authority: what happens to this metric if this change is shown to a random subset of currently eligible users, starting now, for the length of the test, while nobody else changes anything. That sentence has four conditions buried in it, and each one is load-bearing. You can randomize the thing. The unit you randomize is the unit the effect stays inside. The effect appears within the horizon you can afford to wait. And the population you are able to test on is the population you are deciding about. Break one condition and the test does not fail loudly. It runs, it produces a p-value, and it answers a question nobody asked.

This is worth being precise about, because the failure is not statistical. Teams that have internalized power analysis, sequential testing and multiple-comparison corrections still ship the wrong decision, because the estimand — the quantity the design actually estimates — is not the quantity the decision needs. The arithmetic is impeccable and the number is about something else.

Randomizability. You cannot randomize a price change that is announced publicly, a rebrand, a policy that must apply to everyone for legal or contractual reasons, or anything a single competitor will notice and respond to. You also cannot randomize the past, which is where most requests come from.

Non-interference. The standard assumption is that one unit's treatment does not affect another unit's outcome. It is false in every marketplace, every social graph, every shared-inventory system, and every system with a shared budget or a shared queue. If treated riders take the drivers that control riders would have used, the test measures cannibalization, not lift.

Horizon. A two-week test measures a two-week effect. Novelty, primacy, habit formation, churn and reputation all operate on longer clocks, and several of them point the opposite direction from the short-run number.

Population. Tests run on the users you have, during the season you are in, on the surfaces you control. The decision is usually about users you do not have yet, in a season you are not in.

You are probably here because

  • Someone asked what a change did after it was already shipped to a hundred percent
  • A test won by four percent and the company-level metric never moved
  • Two teams ran clean tests on the same surface and both claimed the same revenue
  • The decision is about next year and every test you can run finishes in three weeks

All four are the same failure: the design estimates something real, and it is not the thing the decision needs.

Six questions that break the conditions

These come up constantly. In each case the request sounds like an experiment request and is not one.

1. “What did the thing we already shipped do?” There is no control group, because everyone got it. You are now doing observational work whether you call it that or not. The honest options are a comparison over time against a group that did not receive it, a synthetic control assembled from units that did not, or an interrupted time series with a pre-trend you can show. All three are weaker than a test and all three are far better than the before-and-after difference someone will otherwise put on a slide.

2. “What is the long-run effect?” A recommendation change that increases clicks for three weeks can reduce retention over six months, and the two-week test sees only the first half. Novelty effects decay on a scale of days to weeks; learned behaviour and trust move over quarters. The only clean instrument here is a long-run holdback — a small fraction of users held out for months — and it costs real money in forgone treatment. It is still the cheapest way to find out that a year of wins summed to nothing.

3. “What happens if we roll it out to everyone?” In a two-sided market, a shared inventory pool or a fixed ad budget, the treatment group's gain is partly the control group's loss. A 6% lift measured under 50/50 assignment can be 1% at full rollout, or negative. Switchback designs, cluster randomization by city or region, and budget-split designs each buy back part of the answer at the cost of variance.

4. “Why did it work?” A test gives an average effect, not a mechanism. Splitting that effect into paths — the change improved latency, latency improved conversion — is mediation analysis, and mediation is not identified by the original randomization. The mediator is not randomized. You can test the mechanism directly by randomizing the mediator, which usually means a second experiment, or you can accept that the story is a hypothesis.

5. “What is the effect on the people who use it?” When treatment is offered rather than imposed — an opt-in feature, an email, a beta — the randomized comparison estimates the effect of being offered, which is diluted by everyone who ignored it. The effect on takers is a different quantity, recoverable under assumptions through an instrumental-variables argument that treats assignment as the instrument. It is a real technique with a real assumption: assignment must affect outcomes only through take-up.

6. “Which segments should get it?” A test powered to detect a 2% average effect is not powered to detect segment differences, and slicing it eight ways produces at least one significant subgroup by construction. Heterogeneous-effect estimation with causal forests or an honest split of the data is the right tool, and even then treat the output as a hypothesis for the next test rather than a targeting rule for production.

A test that measures the wrong quantity does not look broken. It looks like a result.

How much a design is worth believing — our working ranking

Randomized test, no interference, outcome inside the horizon
95
Cluster or switchback randomization with a washout window
84
Regression discontinuity at a hard, unmanipulable cutoff
78
Difference-in-differences with visible parallel pre-trends
66
Synthetic control with a long pre-period and placebo tests
60
Matching on observed covariates, no design-based argument
32

Judgment, not measurement. The ordering is the useful part, and it moves with how well the assumption can be checked in your data.

The designs that answer the other questions

Every method below replaces randomization with an assumption. That is the whole trade. The method is easy; defending the assumption is the work, and a team that presents an estimate without naming its assumption has skipped the only part that was hard.

DesignUse whenThe assumption that decides itHow to check it
Difference-in-differencesA change hit some units at a known date and not othersTreated and untreated would have moved in parallel absent the changePlot several pre-periods. Run it on a fake earlier date and get zero
Synthetic controlOne or a few treated units, many untreated, long historyA weighted blend of untreated units tracks the treated pre-periodPlacebo runs on every untreated unit; the real effect should stand out
Interrupted time seriesEveryone was treated at once, and the series is well behavedThe pre-trend would have continued, with seasonality modelledHold out the last pre-period and forecast it; check the residuals
Regression discontinuityTreatment turns on at a threshold on a continuous scoreNobody can precisely manipulate their position at the cutoffDensity test at the cutoff; covariates should be smooth across it
Instrumental variablesTake-up is voluntary but the offer was random or as-good-asThe instrument moves outcomes only through take-upFirst-stage strength, plus a defensible story for exclusion
SwitchbackMarketplace or shared-resource interference at the user levelEffects do not carry over past the washout windowVary the window; the estimate should be stable across widths

Two of these are much stronger than the rest, and it is worth knowing why. A discontinuity design and a well-instrumented offer both have a source of variation you can point at and describe in a sentence: this person fell one point below the threshold, this person was randomly sent the email. Difference-in-differences and matching have no such sentence. They have a model. That is not a reason to avoid them; it is a reason to spend most of the analysis defending them.

Interference, and the design that survives it

Interference is the failure we are called about most, because it is invisible in the test output. Every number looks right. The lift is significant, the sample-ratio check passes, the metric moves. Then the rollout produces a fraction of the predicted gain and nobody can explain the gap.

The mechanism is simple. Where a resource is shared, treatment moves the resource rather than creating it. A ranking change that surfaces a supplier to treated buyers takes that supplier's capacity away from control buyers, so the treatment effect includes an amount that is purely transferred. Under 50/50 assignment the transfer is maximal; at full rollout there is nobody to transfer from.

Three designs push back. Cluster randomization assigns whole cities, regions or supply pools, which contains most spillover inside a cluster at the cost of a much larger effective standard error — the design effect is roughly one plus the intra-cluster correlation times the average cluster size, and with big clusters that factor is not small. Switchbacks randomize time periods within a market, which is efficient when effects settle quickly and biased when they do not, so the washout window is the parameter to sweep. Budget or supply splits partition the shared resource itself, which is the cleanest option and usually the hardest to build.

One diagnostic is cheap and worth running always: if you have any cluster-level variation, estimate the effect at two different treatment shares. If the per-user effect at 50% differs materially from the effect at 10%, interference is present and the naive number will not survive rollout.

Design Note

Keep a permanent long-run holdback, and price it deliberately

A small fraction of users — one percent is a common starting point — held out of a whole class of changes for a quarter or more is the only instrument that measures accumulated effect. It answers the question nobody else can: did a year of individually significant wins add up to anything. The cost is real and computable: the holdback size times your best estimate of cumulative lift. Compute it, write it down, and decide with the number in front of you rather than abandoning the holdback the first time someone asks what it is for.

Defending an assumption is the deliverable

When an estimate rests on parallel trends or on unconfoundedness, the analysis is an argument, and it should be built like one. Four things carry most of the weight.

Show the pre-period, do not summarise it. A plot of both series for six or eight periods before the change tells the reader more than any test statistic. If the lines were not parallel before, say so and stop.

Run the estimator where the answer must be zero. Move the treatment date earlier into a period when nothing happened. Apply the design to an outcome the change could not plausibly affect. A method that finds an effect in both places is measuring something structural about your data.

State how much hidden confounding would overturn it. Sensitivity analysis puts a number on this: how strongly an unobserved factor would have to relate to both treatment and outcome to explain the whole estimate. If the answer is “about as strongly as the strongest variable we do observe,” the finding is fragile and the memo should say it plainly.

Report an interval that includes design uncertainty. A confidence interval from the regression covers sampling noise only. Refit under two or three defensible alternative specifications and report the spread. It is usually wider than the interval, and the gap between the two is the honest measure of how much the method chose the answer.

An observational estimate is an argument with a number attached. Ship the argument, not only the number.

Send us the question and we will tell you what design it needs.

Email the decision you are trying to make, the data you already have, and what has been tried to contact@precisionfederal.com. You get back a short written note naming the estimand, the design we would use, and the assumption it would rest on. One business day. No charge, no meeting, no deck.

contact@precisionfederal.com

Long horizons and the surrogate problem

When the outcome you care about arrives in twelve months and the decision arrives on Thursday, the standard move is a surrogate: an early metric that stands in for the late one. This works only when the surrogate captures the entire path from treatment to outcome, which is a strong condition and almost never checked.

There is a practical way to check it. If you have a library of past experiments where both the short-run and long-run outcomes were eventually observed, you can ask directly how well the short-run reading predicted the long-run one, across those experiments. Teams that do this are frequently unhappy with the answer, and the unhappiness is the value: it is much cheaper to learn that your surrogate has weak predictive power across thirty past tests than to learn it from one year of decisions made on it.

Absent that library, be modest. Use the surrogate as a guardrail rather than a target — ship if the short-run metric does not degrade, rather than shipping because it improved — and keep the holdback that eventually tells you the truth.

When nothing is identified

Sometimes there is no design. The change is one-off, the comparison group does not exist, the history is short, and the honest answer is that the effect is not recoverable from the data you have. This is a legitimate finding and delivering it well is a skill.

Three moves make it useful rather than deflating. Bound it. Even without point identification you can often establish that the effect is between two numbers, and if both numbers point the same way the decision is made. Ask what the answer would change. If the decision is the same at the optimistic and pessimistic ends, further analysis has no value and should not be funded. Design the next one. The most valuable output of an unidentifiable question is a staged rollout for the next change, ordered so that the comparison exists by construction.

That last point is where most of the durable value sits. Rollouts get sequenced for operational convenience — easy regions first, friendly customers first — and that sequencing quietly destroys the comparison. Randomizing the order of a rollout costs almost nothing and converts an unanswerable question into a straightforward one.

Heterogeneity without fooling yourself

“Who should get this?” is a better question than “does this work?” and it is much harder to answer. The average effect is one number estimated from the whole sample. A segment effect is estimated from a slice, so its standard error grows roughly with the square root of how much smaller the slice is, and a test sized to detect a two-percent average effect will not resolve segment differences of similar size. Slicing anyway produces a chart of colourful bars whose ordering is mostly noise.

Three practices make heterogeneity work usable. Pre-declare a small number of segments — three or four, chosen because a mechanism suggests them, written down before the readout. Use an estimator built for the job, such as a causal forest or a doubly robust learner with honest sample splitting, so the same rows that discovered the split do not also estimate its size. Treat the output as a hypothesis: the correct next step after finding a promising segment is a targeted test in that segment, not a targeting rule in production.

There is one exception worth knowing. If the segment split is operationally free and the downside is bounded — for example, suppressing a feature for a group where the point estimate is negative and the cost of being wrong is a small forgone gain — acting on a weak signal is defensible. Write down that you are doing it on a weak signal, and set a date to revisit.

A two-week pass on a causal question

Causal Design Pass

1
Write the estimand in one sentence: whose effect, on what, over what horizon, at what rollout share
Day 1
2
Test the four conditions. Randomizable, non-interfering, inside the horizon, right population
Day 2
3
Pick the design and name its assumption before touching data
Days 3–4
4
Build the pre-period plots, the placebo runs and the density or balance checks
Days 5–8
5
Estimate under three specifications; report the spread, not the best one
Days 9–11
6
Write the memo assumption-first, with the sensitivity number and the decision it changes
Days 12–14

Step one is the one teams skip and the one that saves the project. Half the causal work we are asked to review dissolves at the estimand sentence, because writing it down makes visible that two people in the room wanted two different numbers. The other half survives and gets much easier, because a stated estimand rules out most of the designs immediately.

The mistakes we are called in to fix

  • A before-and-after difference presented as an effect, with no comparison group and no pre-trend shown
  • Segment lifts read off a test powered for the average, then hard-coded into targeting
  • A marketplace test at 50/50 whose lift was mostly transferred from the control side
  • Difference-in-differences with one pre-period, so parallel trends is untestable and unstated
  • A surrogate metric never validated against the long-run outcome it stands in for
  • Intent-to-treat reported as the effect on users, understating it by the take-up rate
  • Rollout order chosen for convenience, destroying a comparison that was free five minutes earlier
  • A confidence interval reported as total uncertainty, with three other specifications quietly discarded

Before you call a number causal

  • The estimand is written down: whose effect, on what, over what horizon, at what share
  • The source of variation is stated in one sentence a non-specialist understands
  • The identifying assumption is named, and someone has tried to break it
  • Pre-period behaviour is plotted, not summarised
  • A placebo run on a date or outcome with no effect returns approximately zero
  • Interference has been considered explicitly and ruled in or out
  • The reported interval reflects specification spread, not just sampling noise
  • Sensitivity is quantified: how much hidden confounding would overturn the result
  • The memo states which decision changes at which value
  • If nothing is identified, that is the finding, with a design for next time

Bottom line

Randomization is the strongest tool available and its reach is narrower than most roadmaps assume. The discipline that matters is not statistical sophistication; it is writing the estimand down first and checking it against four conditions before anybody opens a notebook. Where the conditions hold, run the test and trust it. Where they fail, pick a design whose assumption you can state and attack in public, spend the analysis defending that assumption, and report an interval wide enough to be honest. And where the question is genuinely unanswerable, say so and use the hour to make the next question answerable instead. That last habit compounds faster than any method on this page.

Common objections

We have a lot of data. Does that not solve the identification problem?

No. Sample size shrinks the confidence interval around whatever quantity the design estimates; it does nothing to move that quantity toward the one you wanted. A biased estimate computed on ten billion rows is a precise wrong answer, and the narrow interval makes it more persuasive rather than less. Volume helps with power and with heterogeneity. Identification is a separate axis.

Can a machine learning model just predict the counterfactual?

Only under the same assumption every observational method needs: that the variables you conditioned on account for the reasons treatment was assigned. Modern estimators handle the nuisance functions better than a linear regression does, which is a real gain, and they do not create identification. If an unmeasured factor drove both the treatment and the outcome, a stronger predictor fits the confounding more faithfully.

Is a holdback worth its cost when every test wins?

That is precisely when it is worth the most. A sequence of individually significant short-run wins that sum to no change in the long-run metric is a common and expensive pattern, and only a long-horizon holdout can detect it. Price the holdback explicitly — its size multiplied by your best estimate of cumulative lift — and compare that number against the cost of a year of decisions made on an unvalidated surrogate.

The change is going out regardless. Why bother with the design?

Because the design costs almost nothing at that point and buys the ability to answer questions later. Randomize the order of a staged rollout, keep the last wave a month behind, and record who received what and when. None of this delays the launch, and it converts a permanently unanswerable question into an ordinary one.

Frequently asked questions

How do you measure the effect of something already rolled out to everyone?

By finding variation that survived the rollout: units that received it later, a comparable population that never did, or a threshold that decided who got it first. Those give you difference-in-differences, synthetic control or a discontinuity design. If none of the three exists, an interrupted time series with a modelled seasonal component is the remaining option, and it is weak enough that the memo should say so.

Why do marketplace tests overstate the effect of a rollout?

Because treated and control users compete for the same supply. Part of the measured lift is capacity moved from control rather than created, and at full rollout there is no control side to move it from. Cluster or switchback randomization contains the spillover; comparing the estimated effect at two different treatment shares tells you quickly whether it is present.

Is difference-in-differences trustworthy?

It is trustworthy in proportion to how convincingly you can show parallel pre-trends and how well a placebo run behaves. With several clean pre-periods and a placebo that returns zero, it is a reasonable basis for a decision. With one pre-period it is an assumption nobody can inspect, and it should be labelled that way in the memo.

How long should a long-run holdback run?

Long enough to cover the mechanism you are worried about. Novelty decays in days to weeks; habit, trust and churn move over one to two quarters. A common structure is a permanent one-percent holdout read quarterly, with the forgone-lift cost computed and reviewed rather than assumed away.

When is a subgroup result from a test safe to act on?

When the subgroup was declared before the test ran, the test was powered for it, and the direction is consistent with a mechanism you can state. Everything else is a hypothesis: estimate heterogeneity with an honest method, then confirm the targeting rule in a follow-up test rather than shipping it from the slice.

1 business day response

Have a question randomization cannot reach?

Send the decision, the data you hold and what has already been tried. Our engineers will come back with the estimand, the design that fits it, the assumption it rests on and the checks we would run — or take the analysis and build it as a scoped piece of work. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
Causal InferenceExperimentationProduct AnalyticsDecision Support