You are buying a panel, not a measurement
Whatever the deck says, a third-party dataset is almost never a measurement of the world. It is a sample of the world produced by some mechanism — a set of cards, a set of phones, a set of scraped pages, a set of satellite passes — and everything you can learn from it is bounded by how that sample was assembled and how it changes. The single most useful hour in an evaluation is the one spent establishing what the panel is, because it determines which questions the data can answer and which it will answer confidently and wrongly.

This applies well beyond market data. A team buying a labelled corpus for model training is buying a panel too: some annotator pool, some sampling rule, some drift. The questions below transfer with the nouns changed.
You are probably here because
- A vendor sent a backtest and you cannot tell whether it means anything
- A trial ended and the team is arguing about what it showed
- A dataset that worked two years ago stopped working and nobody knows when
- Renewal is in six weeks and there is no written basis for the decision
The first is a multiple-testing and point-in-time question. The second is usually a trial that was never designed to be conclusive. The third is panel drift, and it is measurable if you kept the deliveries.
The five questions before any trial
Ask these before requesting a sample. Two of them frequently end the conversation, which is a good outcome in week one and an expensive one in month nine.
How long is the history, and is it point-in-time? A five-year history assembled last month from current records is not five years of evidence. What matters is whether each record carries the timestamp at which it was first available, and whether the vendor kept its own past deliveries. Ask directly: were these files reconstructed, or archived as they were sent? The answers differ more often than the marketing suggests.
How was the panel built, and how does it churn? Where do the contributors come from, what is the acquisition channel, what fraction leaves each month, and has the mix changed? A panel that grew through one channel until 2023 and another after it is two datasets with one name, and any level comparison across that boundary is meaningless while the growth rates may still be fine.
What is the delivery latency distribution? Not the average. The median, the ninetieth percentile, the worst case, and the failure rate. A feed that arrives on day two for eighty percent of records and day nine for the rest is a day-nine feed for any use that needs completeness.
How do I get from their identifiers to mine? This is the question that decides project cost, and vendors under-describe it consistently. They have a merchant string, a domain, a place name, a vessel identifier. You need an entity you can act on. If the mapping is not supplied, and often it is not, the mapping is your project.
What may we do with it? Derived-data rights, redistribution across internal teams, what survives termination, whether the panel's own permissions cover your use. These constrain architecture, not just legal review, and finding out afterwards is how a platform gets rebuilt.
Coverage is a function, not a number
“We cover eighty percent of the market” is not a specification. Coverage varies by segment, by geography, by size of the thing observed, and by period, and the variation is usually the interesting part. What you want is a coverage function you estimated yourself, not one the vendor asserted.
The estimate is straightforward wherever some subset of your targets publishes a ground-truth figure. Take the entities that report, regress the panel's measure against the reported measure, and look at the relationship across segments and across time. Three things fall out: whether the panel tracks reality at all, where it tracks better and worse, and — most valuable — whether the relationship is stable. A coefficient that drifts steadily over three years tells you the panel is changing underneath a signal that appears constant.
Do this per period rather than pooled. A pooled fit averages away exactly the instability you are trying to detect, and instability is the thing that ends most datasets: not that the signal was never there, but that the panel it depended on became a different panel.
Point-in-time, or it did not happen
Every record needs the timestamp at which it became available to you, and it must be the vendor's publication time rather than the event time. Without it, any historical evaluation reads the future by however long the reporting lag is, and reporting lags are neither short nor constant.
When a vendor cannot provide it — a common and not always disqualifying answer — you have two options. Build it yourself by capturing deliveries as they arrive, timestamping them, and evaluating only on the window you captured. That is a legitimate reason to run a trial slowly: six months of your own captured deliveries beats five years of reconstructed history for deciding whether something is real. Or accept the reconstructed history for exploration only, with a written note that no result from it may support a purchase.
| Data type | The failure mode that gets people | Cheapest early check |
|---|---|---|
| Consumer transaction panels | Panel composition drifts; level comparisons break across the drift | Coverage regression re-estimated per quarter |
| Web and app telemetry | Measurement changes when the observed party changes their site or app | Diff the collection schema against last quarter's |
| Scraped listings and prices | Silent collection gaps look like real declines | Plot record counts per source per day and look for cliffs |
| Satellite and imagery-derived counts | Weather and revisit rates make missingness correlate with the thing measured | Check whether gaps cluster by season and location |
| Shipping, customs and logistics records | Reporting lags vary by jurisdiction and are treated as constant | Lag distribution by origin, not pooled |
| Labelled corpora for training | Annotator pool changed mid-collection; label meaning shifted with it | Re-label a stratified sample and measure agreement by batch date |
How many independent observations do you actually have?
This is where most trials quietly become uninformative. A dataset covering three hundred entities quarterly for five years advertises six thousand observations. It does not have six thousand independent ones. The entities move together, so the cross-section contributes far less than its count suggests, and the honest denominator is closer to the twenty quarters than to the six thousand rows.
Twenty is a small number. It will not distinguish a modest true effect from noise, and no amount of modelling repairs that. Working this out in week one changes the trial: you stop trying to prove a small effect you cannot resolve, and you spend the time on questions the sample size can actually answer, such as whether coverage is stable and whether the data agrees with a known quantity.
Why trials fail to be conclusive — what we see, roughly ranked
Our ranking from evaluations we have worked on, not a survey. The ordering is the useful part, and the top two are decided before any modelling starts.
Write the criteria down before you look
If you try forty transformations of a dataset against a target, some of them will look good. At a conventional threshold you would expect two false positives from forty independent tries, and the transformations are not independent, which makes it worse rather than better. The result you will present is the maximum over the search, and the maximum over a search is not an estimate of anything.
Three practices fix most of this without heavy machinery. Write down the transformations and the success threshold before the data arrives, and count every variation you subsequently try. Hold out time, not rows — the last twelve months, untouched, opened once. And report the whole distribution of what you tried alongside the best result, because a top result that sits at the edge of a broad distribution of failures means something very different from one that sits among several near-misses.
Does it add anything to what you already have?
A dataset that predicts your target is not necessarily worth buying. The question is whether it predicts anything your existing inputs do not. The test is mechanical: fit your current inputs first, take the residual, and evaluate the candidate against that. Vendors do not do this and cannot, because they do not know what you already run.
We have watched a genuinely good dataset fail this test — real signal, honestly built, almost entirely explained by two things already in the stack. That is a fair no, and it is a much better outcome than a yes that adds a subscription, an ingestion pipeline, a licensing obligation and an ongoing reconciliation for a marginal contribution nobody can find a year later.
Ask how many clients have this, and ask what happened to the ones who left
Any edge available to everyone is priced in by everyone. A vendor with three hundred subscribers is selling a utility, which may be exactly what you want and should be priced like one. A vendor with four is selling something else, and exclusivity or a limited-subscriber term is negotiable more often than buyers assume. The churn question is the more revealing half: a vendor who can describe why customers left, specifically, is usually the one telling the truth about the rest.
Send us the sample and the vendor's backtest.
Email a data dictionary, a representative extract and whatever performance claim came with it to contact@precisionfederal.com. You get back a written note on what the trial can and cannot establish, the checks we would run first, and the questions we would put to the vendor. One business day. No charge, no meeting, no deck.
contact@precisionfederal.comThe mapping is usually the project
Vendors deliver their own identifiers. You need yours. Between the two sits a matching problem with no shared key, a name field written by whoever typed it, and a mapping that has to be maintained rather than built once, because the source keeps changing underneath it.
Two things follow. Scope the mapping work into the trial rather than after it, because a trial that runs out of time at eighty percent mapped produces a result about the easy eighty percent — and the easy ones are systematically the large, well-known entities, which is a bias in the direction of optimism. And keep the mapping as its own versioned asset with match evidence attached, so that when a number is disputed you can show why two records were linked instead of asserting that they were.
What a fair trial looks like
Eight-week evaluation
Two disciplines make this work. Capture every delivery you receive during the trial, with your own arrival timestamps, and never overwrite one; that archive is a point-in-time record you will not be able to reconstruct afterwards, and it keeps its value even if you decline. And keep an evaluation log listing every transformation tried, including the ones abandoned after ten minutes, because the count is what lets anyone interpret the final number.
Price it against the alternative, including doing nothing
Quoted prices for a single dataset span an enormous range, from the low tens of thousands a year to figures that need a committee, and the spread within a category is wide enough that a quote tells you little on its own. What you can do is bound the value: estimate what the decision improves by, honestly, and compare against the total cost, which is the subscription plus the ingestion build plus the ongoing mapping maintenance plus the licensing constraints on how you may use the output.
That last term is the one left out of business cases. A dataset that cannot cross an internal boundary, or whose derived outputs cannot be shown to a client, is worth substantially less than the same data without those restrictions, and the difference belongs in the number rather than in a footnote.
The mistakes we are called in to fix
- Evaluating on a reconstructed history and treating the result as evidence of what was knowable
- Reporting the best of forty transformations without reporting that there were forty
- Pooling the coverage estimate, averaging away the drift that ends the dataset
- Judging latency by the mean when the tail decides whether it is usable
- Leaving the mapping until after the trial, so the result describes the easy entities only
- Never testing against existing inputs, and buying a restatement of what you own
- Accepting a curated sample as representative of the delivered feed
- Signing before reading the derived-data terms, then designing around them afterwards
Before you sign
- Publication timestamps on every record, or your own captured delivery archive
- Panel construction and churn described in writing, with the change points named
- Coverage estimated by you, per segment and per period
- Latency measured as a distribution, including the failure rate
- Mapping built and its coverage measured, not assumed
- Success criteria and transformation list written before the data arrived
- An untouched holdout period, opened exactly once
- Evaluation performed on the residual after existing inputs
- Derived-data, redistribution and post-termination terms read by an engineer
- A written memo including everything that did not work
Bottom line
Most of what decides an alternative data purchase is settled before any modelling: whether the history reflects what was knowable at the time, whether the panel is stable enough for a comparison to mean anything, whether you can map their identifiers to yours, and whether the thing adds anything beyond what you already own. Those four are cheap to check and they eliminate most candidates in a fortnight. Spend the trial establishing them, keep an honest log of what you tried, and treat a well-supported no as a successful outcome — it is the cheapest one available after the money is spent.
Frequently asked questions
Long enough to capture your own deliveries, which usually means months rather than weeks if the vendor cannot supply publication timestamps. Eight weeks is workable when the history is genuinely point-in-time and the mapping is supplied. When neither holds, a short trial mostly measures how quickly your team can build a mapping.
Read it for what it tells you about the data rather than about the performance. The useful questions are which universe it covers, whether the history was archived or reconstructed, how many variants were tried, and whether the entities that disappeared are still in the sample. A vendor who answers those precisely is worth more attention than one with a better chart.
The effective sample was too small to resolve the effect being looked for, and nobody worked that out at the start. Entities move together, so a cross-section of hundreds contributes far fewer independent observations than its row count suggests. Establishing the honest denominator in week one changes which questions the trial asks.
Re-estimate the coverage relationship period by period against a ground-truth series. A stable relationship with a weakening signal suggests the information is being priced in or the world changed. A drifting relationship means the panel changed and the signal was never comparable across the boundary in the first place.
Nearly all of it. A labelled corpus is a panel with an annotator pool and a sampling rule, it drifts when either changes, and it needs the same questions about provenance, licensing and what may be done with derived outputs. Substitute label agreement for coverage and the method transfers directly.
