Why three tools give you three answers
The meeting goes the same way in every company. The ad platform reports four hundred conversions. The other platform reports three hundred and fifty. The email tool claims two hundred. The warehouse says there were five hundred and twenty orders. Somebody says the tracking is broken. The tracking is not broken. Each platform is reporting the conversions it can see and claiming credit for every one it touched, and the same order is being counted by three of them.

Across most mid-size advertisers the sum of platform-reported conversions runs somewhere between twenty and sixty percent above the order count in the system of record. That gap is not fraud and it is not a bug. It is the predictable result of every seller marking its own homework inside its own attribution window, with its own view-through rules, on a population it cannot fully observe.
The second thing to accept is that the click-path data got worse and is not coming back. Mobile tracking permission changed what an app can observe. Browser storage lifetimes shortened. Walled gardens report aggregates instead of user-level events. Meanwhile the parts of demand that were never trackable — a colleague's recommendation, a podcast mention, a link pasted in a group chat — are a larger share of how people find things than they were a decade ago. Any measurement plan whose foundation is a complete user journey is being built on a surface that is still eroding.
You are probably here because
- Your agency's report and your finance team's report disagree, and both are internally consistent
- Branded search shows a spectacular return and you suspect it is buying people who already knew you
- Someone wants to shift a seven-figure budget on the strength of a dashboard nobody can explain
- You paused a channel for two weeks, revenue did not move, and nobody could say what that proved
The last one is closest to a real answer. It just needs to be designed rather than stumbled into.
Separate the two questions
Attribution assigns credit for orders you already have. It is a bookkeeping exercise over observed touchpoints, and every rule for doing it — last touch, first touch, linear, time decay, a trained multi-touch model — is a choice about how to split a fixed pie. It is useful for tactical routing: which creative to rotate off, which keyword to pause, which audience to expand. It is fast, granular and directional.
Incrementality asks what would have happened without the spend. That is a counterfactual, and no amount of observational click data answers it. Only an experiment, or a model that stands in for one, can. This is the question behind every budget decision, and it is the question a last-touch dashboard silently pretends to answer.
Most attribution fights are two people answering different questions and thinking they disagree. Say out loud which one is on the table before anyone opens a spreadsheet.
The plumbing has to exist first, and it takes longer than the modeling
Everything below needs one table: spend by channel by day, joined to revenue by day, on a shared calendar. It sounds like an afternoon. Plan for four to eight weeks, because of the things that are always wrong.
Time zones and reporting days do not line up between platforms and your warehouse, which puts a systematic shift into a series where a one-day lag matters. Currency conversion is applied at different points. Platform spend excludes agency fees and production, so the return you compute is not the return the business experiences. Refunds, cancellations and returns arrive weeks later and belong back on the original day, not the day the credit was issued. Discount codes double as attribution and as margin reductions and are usually recorded once.
And the lag. For a considered purchase — anything with a multi-week sales cycle — spend in March produces revenue in June. Regress June revenue on June spend and the model learns nothing except that both grew. You either model the lag explicitly or you work at a grain coarse enough that it washes out, and pretending the problem away is the most common reason a mix model produces confident nonsense.
Experiments are the only ground truth
A geographic holdout is the workhorse. Split markets into a test group and a control group, turn a channel off in control for a defined window, and compare the difference between groups against the difference before the test. The design is old, it is well understood, and it is the only evidence in this article that does not depend on an assumption someone can dispute.
The part that gets skipped is the power calculation, and skipping it is how companies run a six-week test that could never have detected the effect they were looking for. Rough shape of the arithmetic: if a channel drives a three percent lift on total revenue and weekly revenue in a market swings eight percent for reasons unrelated to marketing, you need many markets and many weeks to see three percent through that noise. Run the numbers before the test, and if the answer is that you cannot detect the effect you care about, that is a finding. It means the channel is small enough that the decision does not need this level of proof.
| Design | What it answers | What it costs | When it is the right tool |
|---|---|---|---|
| Geographic holdout Channel dark in matched markets | Incremental revenue for a whole channel | Real forgone revenue in control markets for 4–8 weeks | National or multi-market spend with 20+ comparable geographies per arm |
| Platform-run lift test Randomised holdout inside the ad system | Incremental conversions for a campaign | Little or nothing beyond minimum spend | Quick reads, with the caveat that the seller is scoring the test |
| Branded search pause | How much paid brand traffic organic would have captured anyway | A few weeks of nerves | Almost always worth running once a year |
| Budget-split test Same channel at two spend levels by region | The shape of the returns curve, not just whether it works | Longer window, more markets | Deciding how much more to spend, not whether to spend |
| Creative or landing A/B | Relative performance inside a channel | Traffic and time | Optimisation. It cannot tell you the channel's incremental value |
Run branded search first. It is cheap, it is uncomfortable, and it is where the largest single correction usually lives. In a typical account a meaningful share of paid clicks on your own brand terms would have arrived through organic results at no cost, and estimates in the range of half to nine tenths are common enough that it is worth finding your own number rather than arguing about someone else's.
How much a method should move a budget decision — our weighting
How much weight we give each source when the question is where the next dollar goes. Judgment, not measurement — the ordering is the part worth keeping.
The survey line surprises people. A single open question at checkout is imprecise, biased toward the most recent memory, and still the only instrument that sees the podcast, the friend and the conference. It is worth its place as a cross-check, and worth nothing as a primary number.
Media mix modeling, and when it cannot work
A mix model regresses a business outcome on spend across channels plus seasonality, price, promotions, distribution and whatever else moves the business, with two transformations that make it a marketing model rather than a plain regression. Adstock captures the fact that spend today produces sales over the following days or weeks. A saturation curve captures diminishing returns, so the model does not conclude that ten times the budget produces ten times the sales.
It needs two things, and the second is where projects die. It needs history: two to three years of weekly data, meaning roughly a hundred to a hundred and fifty observations, against which you are estimating a lot of parameters. And it needs variation. If you have spent the same forty thousand dollars a month on every channel for two years, and all of them move together with the season, the model cannot separate them. There is no statistical technique that recovers a signal the data does not contain. It will still return coefficients, with confident-looking numbers, and they will be arbitrary.
That has an actionable consequence. If you intend to build a mix model next year, start creating variation this year: stagger channel flighting, run regional pulses, hold a channel out for a month at a time. Deliberate variation is what makes the history informative, and it is nearly free compared with the model.
When it does work, the useful output is not a point estimate of return. It is a saturation curve with an uncertainty band, which answers the actual question — what happens to revenue if this channel gets another two hundred thousand dollars — and it makes the honest answer visible when the band is wide.
A return of 2.4 with an interval of 0.9 to 4.1 is not a number to reallocate against
Point estimates travel. Intervals do not, unless you force them to. Put the interval in the headline of the slide, not in an appendix, and put the actions in bands: channels we are confident about, channels we are not, and channels we cannot see at all. A leader who is told that three channels are measured and two are guesses will make a better decision than one handed five numbers of equal-looking precision.
Calibrating the model to the experiments
The two methods are usually presented as rivals. They work far better joined. Run experiments on the two or three channels you can test. Use those results as priors in the mix model, so the model is anchored to measured causal effects where you have them and interpolates where you do not. Then check the model's prediction for a channel you did not use in calibration against a later test on that channel. If it holds, the model has earned some trust. If it misses badly, you have learned that before betting a budget on it.
This is also the cadence that works organisationally. Two or three experiments a year, a mix model refreshed quarterly and recalibrated when new experimental results land, and platform data used only for in-channel routing. Anything faster is measuring noise and anything slower is not usable for planning.
The problem is often governance, not statistics
An agency compensated on the performance it reports is being asked to score its own work. That is not a comment on anyone's integrity; it is a structure that reliably produces optimistic numbers everywhere it exists. The same applies to an internal team whose bonus depends on a channel's reported return.
The fix is boring and it works. One number, produced by a group that does not own the spend, from a definition written down in advance, with the methodology change log visible. Define a conversion once. Define the attribution window once. Decide before the quarter which source is the one that counts. Most of the reconciliation pain in this discipline comes from a definition that quietly moved.
When you do not need any of this
Under roughly a million to two million dollars a year in working media, a full measurement stack costs more than the decisions it improves. What you need instead is a weekly spreadsheet with spend and revenue, a clean join to orders, a post-purchase survey question, and two well-designed holdout tests a year. That is a few days of work and a few weeks of patience, and it will beat a purchased attribution platform that reconciles nothing.
You also do not need a mix model if you only have one meaningful channel. Turn it off in half your markets for six weeks. You will learn more than a model would tell you, and you will have learned it from your own business.
What a defensible measurement pack contains
- One revenue definition, written down, net of refunds, matched to the finance number
- Total spend including fees and production, not the platform's media figure
- A named experiment per major channel with its date, design and power calculation attached
- Every reported effect carries an interval, and channels with no evidence are labelled as such
- The lag structure stated explicitly for any business with a sales cycle beyond a week
- A change log of methodology, because a number that moved for a definitional reason is not a result
- A standing calendar of variation — pulses and holdouts scheduled ahead, not improvised
The mistakes we get called about
- Summing platform-reported conversions and treating the total as the truth
- A mix model built on two years of flat, perfectly correlated spend, which cannot identify anything
- A holdout test with six markets per arm, powered to detect an effect twice the size of the real one
- Regressing this month's revenue on this month's spend in a business with a ninety-day sales cycle
- Branded search never tested, and carrying the highest reported return in the account
- Point estimates in the board deck and the confidence intervals in a hidden slide
- Attribution window changed mid-quarter, then the change compared against the old series
- A multi-touch model rebuilt every year on click data that keeps getting thinner
Bottom line
Pick the question first. If it is which creative to cut, the platform data is fine and fast. If it is where the next two hundred thousand dollars goes, only an experiment or a model calibrated to experiments can answer it, and both need plumbing that takes longer than the analysis. Run a branded search holdout this quarter, start creating deliberate variation in the spend now so a mix model has something to learn from later, and get the reporting out of the hands of the people being graded by it. Report intervals. Say plainly which channels you cannot see. A leader who knows which two numbers are guesses makes better decisions than one holding five numbers that all look equally solid.
Frequently asked questions
Because each platform claims every conversion it touched inside its own window, using its own view-through and click rules, and the same order is counted by several of them. It is double counting rather than error. Compare each platform against your own order table, expect the sum to run twenty to sixty percent high, and treat the warehouse as the denominator.
Four to eight weeks is typical, but the honest answer comes from a power calculation on your own data. Take the week-to-week variation in revenue per market, the number of markets you can assign to each arm, and the smallest lift you would act on, and compute how long it takes to see that lift through the noise. If the answer is a year, the channel is not big enough to warrant the test.
Two to three years of weekly observations is the usual floor, and the amount of variation in the spend matters more than the number of rows. Flat budgets across channels that all move with the season leave nothing to identify. If your history looks like that, spend the next few quarters deliberately varying spend by region and by flight before commissioning a model.
Not dead, but demoted. It is still useful for routing decisions inside a channel where you can observe the clicks. As a basis for allocating budget across channels it was always resting on complete journey data, and that data has been thinning for years. Use it tactically and do not let it near a budget meeting on its own.
Yes, and it is the cheapest informative test available. Pause paid brand terms in a set of markets for a few weeks and watch how much of the traffic organic recovers. The cannibalisation share varies widely by category and by how crowded your brand terms are, which is exactly why the number has to come from your own account rather than from a published average.
