First find out whose defect it is
Every serious data organization has a list of grievances against its main vendor. Coverage gaps, late files, values that changed without explanation, a support queue that answers in a week. The list is usually real. It is also, in our experience, a mixture — and a meaningful share of the items on it turn out to originate on the buyer’s side, in the mapping and matching layer between what the vendor sent and what the business uses. That is not a criticism of anyone. It is a consequence of the fact that nobody instruments the boundary.
The pattern repeats. An analyst reports that the vendor is missing coverage on a set of entities. The entities are in the file. They failed to match because the internal identifier mapping was built years ago against a different vendor key and has been drifting since. Or a field is reported as wrong; the field is right, and a transformation two steps downstream is applying an outdated code list. Or the data is reported as late; it arrived on time and the load job was queued behind a backfill.

So the first step is never to shop. It is to take the complaint list and, for each item, establish where the defect enters. That means comparing the raw delivered file against what your consumers see, which requires that you kept the raw delivered file — and if you did not, that is the first thing to fix regardless of what you decide about the vendor. Two or three weeks of that work reliably reorganizes the list, and it produces the only artifact that makes the rest of the decision honest: a defect log with an attribution column.
You are probably here because
- A renewal is coming and somebody asked whether you should still be paying it
- An incident reached a client and the vendor’s explanation did not satisfy anyone
- A competitor product is being pitched to your team and it demonstrates well
- Coverage you need has been on the roadmap for two years
The measurement section is what turns any of those into a decision you can defend. The switching-cost section is the part that is always underestimated. The table of structural reasons is the short list that actually justifies a move.
Measure against your universe, not against a demo
A vendor evaluation built on the vendor’s own materials measures their marketing. An evaluation built on your universe, over a period you lived through, measures the thing you are buying. Five dimensions cover most of it, and all five can be computed from data you already have.
Coverage on the set you care about. Not the total entity count, which is a headline number and nearly meaningless for comparison. The share of your universe present, computed by a matching procedure you have validated, with the unmatched residual inspected by a human rather than assumed to be genuinely absent.
Timeliness against your window. The distribution of arrival times relative to when you need the data, over months. The mean is uninformative; the tail is the whole story, because your process breaks on the late days and nothing else.
Correctness where you can adjudicate it. Pick fields where an independent source of truth exists — a public filing, a second vendor, a customer-confirmed value — and sample. This is slow, and it is the only dimension where an outside opinion can be replaced by evidence, so it is worth the effort on the two or three fields that matter most.
Correction and restatement behavior. How often values change after publication, how far back, and whether you were told. A vendor that corrects openly is better than one whose numbers never move, because numbers that never move in a domain that does are not a sign of accuracy.
Resolution time on your disputes. Measure it from your own ticket history: how long from report to confirmed fix, and what share were closed without one. This is often the honest reason people want to switch, and it deserves to be stated as a number rather than a feeling.
Reasons that justify a replacement, and reasons that look like they do
Once the measurement exists, most complaint lists sort cleanly into two groups. One group describes something structural about the vendor — a property of their business, their sourcing or their engineering that will not change because you asked. The other describes a bad period, a bad relationship, or a problem you own.
| Signal | Structural? | What it means |
|---|---|---|
| They do not collect what you need | Yes | A coverage gap in their sourcing, not their processing. Roadmap promises to close it are worth what the last three roadmap promises were worth |
| No effective dates, no history, no deltas | Yes | A data model limitation that caps what you can ever build on top, regardless of how good the values are |
| Repeated silent structural change | Yes | Fields retyped or renamed without notice, more than once, after being raised. This is an engineering-culture fact and it will keep costing you incidents |
| Restatements you learn about from clients | Yes | Their release process does not treat you as a consumer who needs notice. Rarely fixed by escalation |
| Terms that trap your history | Yes | If you cannot retain what you already received after termination, the relationship gets more expensive to leave every year you stay |
| A bad quarter of timeliness | Usually not | Ask what changed. Vendors have incidents, migrations and staff turnover like everyone else. The question is whether the trend recovers |
| Price increase at renewal | No | A commercial question. Real, sometimes decisive, but it is negotiated rather than escaped, and switching to escape a price rise usually costs more than the rise |
| A rival product demonstrates better | No | Demonstrations are built on the cases that demonstrate well. Nothing counts until it has run against your universe for a full period |
| Your team dislikes the account manager | No | Worth fixing, and fixable. Ask for a different one before you consider a migration |
The switching costs nobody counts
The subscription price difference is the easiest number to compute and usually the smallest term in the equation. What actually determines whether a switch is worth it is a set of costs that live in engineering time and in the comparability of your own published output.
Identifier remapping. Every internal key mapped to the old vendor’s key must be remapped to the new one, and the two vendors will not agree on entity boundaries. One will treat a parent and its subsidiary as one entity, the other as two. The residual that cannot be mapped automatically has to be worked by a person who understands the domain, and that residual is almost always larger than the initial estimate.
History and comparability. The new vendor’s history is not the old vendor’s history. Values will differ for real methodological reasons, and every series that crosses the switch date has a discontinuity in it. If you publish anything to clients, that discontinuity is now something you have to explain, document, and possibly restate. In our experience this is the cost that most often turns a switch decision around when it is finally quantified, and it is almost never in the first version of the business case.
Everything downstream that was tuned. Models fitted on the old distributions, thresholds set against the old scales, quality rules calibrated against the old defect patterns, dashboards with the old field names, documentation, and a support team who know which oddities are normal. All of it needs revisiting, and much of it is undocumented knowledge in people’s heads.
The parallel period. You will pay both vendors at once for as long as the comparison and cutover take. Plan for that period to be longer than the plan says, and treat any proposal that omits it as incomplete.
Where the effort in a vendor switch actually goes — our read
Relative weight in the switches we have worked on, as a judgment rather than a measurement. The bottom row is the one migration plans are usually built around.
Run the comparison so the answer means something
A trial that consists of the challenger sending a sample file and someone eyeballing it proves nothing. A comparison worth acting on has four properties, and they are not expensive to arrange.
It runs on your universe over a period you already lived through, so you know what happened and can adjudicate disagreements against reality rather than against each other. It runs long enough to include a month-end, a quarter-end and at least one holiday, because that is when both vendors’ operational behavior is visible. The matching procedure is validated before any comparison runs, because a coverage difference that is really a matching difference will otherwise decide your evaluation for you. And disagreements are adjudicated against a third source where one exists — the point is not which vendor differs from the other, but which one is right when they differ, and you cannot learn that from the two of them alone.
One more thing worth doing, and rarely done: send both vendors the same set of genuine defects you have found, through their normal support path, and measure what comes back. The response tells you more about the next five years than the data comparison does. A vendor who confirms, explains, dates the fix and follows up is describing how they will behave when the problem is expensive. So is one who does not.
Consider the two options that are not replacement
Second-sourcing. For the fields that matter most, taking a limited feed from a second provider gives you a continuous adjudication signal, a fallback when the primary has an incident, and an accurate picture of relative quality that no evaluation period can match. It costs more than one subscription and less than a migration, and it converts the vendor question from a periodic crisis into an ongoing measurement. For most organizations that publish anything, this is the answer more often than replacement is.
Renegotiating from a measured position. A defect log with attribution, a timeliness distribution and a resolution-time series change a renewal conversation completely, because they replace assertion with evidence. Vendors respond to specifics. Ask for what you actually need — effective dates on every record, notice before structural change, a named engineering contact for defects, the right to retain delivered history after termination — rather than a discount, which is the ask that is easiest to grant and changes nothing about your problem.
That last item deserves a sentence of its own. The right to keep and continue using the data you already received, after the relationship ends, is the term that determines how expensive leaving will be. It is far easier to obtain at renewal, when you are staying, than at termination, when you are not. Read what your current agreement says about it before you need to know, and if the answer is unclear, that is worth resolving in writing this year rather than next.
Common mistakes
- Shopping before measuring, which guarantees the new vendor is evaluated on their materials and the old one on your grievances
- No attribution on the defect log, so problems that originate in your own mapping layer are counted against the vendor
- Comparing coverage counts instead of coverage on your universe under a validated match
- A trial too short to include a quarter-end, which is exactly when operational differences appear
- Omitting the parallel period and the remapping effort from the business case
- Ignoring the discontinuity in published series until a client asks about it
- Discovering exit terms at termination, when there is nothing left to negotiate with
- Switching to escape a price increase, which usually spends more than it saves
Where you do not need us
Most of this is work your own team is better placed to do than any outsider, because it depends on knowing which fields matter and which oddities are normal. The defect log with an attribution column, the timeliness distribution from your own load logs, the resolution-time series from your own tickets — those are queries against data you already have, they take days rather than months, and they are the entire foundation of the decision. Nobody should be paid to produce them for you.
The parts where outside help earns its place are the ones that need engineering capacity you cannot spare while also running the business: building a matching harness good enough that a coverage comparison is trustworthy, executing the remapping residual, and running a parallel reconciliation that produces a defensible answer rather than two tables that disagree. And if the honest conclusion is that the incumbent is adequate and the problem is your own boundary layer, that is a good outcome, it is cheaper than a migration, and we would rather say so than sell the larger project.
Bottom line
Replace a data vendor when the reason is structural — they do not collect what you need, their model cannot represent history, or they change things underneath you without notice — and not when the reason is a bad quarter, a price rise or a good demonstration. Establish which it is by measuring against your own universe with attribution on every defect, then price the switch including remapping, history discontinuity and the parallel period rather than the subscription difference alone. Before you commit, look hard at second-sourcing the few fields that matter, because it answers the underlying question continuously and costs a fraction of a migration. And whatever you decide, secure the right to keep what you have already received, at the renewal where you are still staying.
Frequently asked questions
Long enough to include at least one month-end, one quarter-end and one holiday period, because that is when the operational differences between two vendors become visible. A shorter window measures steady-state behavior, which is where vendors most resemble each other. The exact length depends on your calendar, but a single month is rarely enough to learn anything decisive.
Often neither, in the sense you mean — they are frequently measuring slightly different things under different methodologies, and both are internally consistent. Adjudicate a sample against a genuinely independent source before concluding anyone is wrong. Then decide separately what to do about the discontinuity in any series you publish, because that is a disclosure question rather than a data question.
For the small number of fields that carry the most consequence, frequently yes. You get continuous adjudication, an immediate fallback during an incident, and an ongoing measurement of relative quality that no evaluation period produces. Scope it to those fields rather than the full universe — a limited second feed and a full second subscription are very different commitments.
Only if your agreement says so, and the answer varies enormously between vendors and between contracts with the same vendor. Read the termination and post-termination use sections of your current agreement before assuming either way, and take the question to counsel rather than to the account team. If the answer is restrictive, the time to change it is a renewal where you intend to stay.
That is a good result. It is cheaper to fix than a migration, it is entirely within your control, and it would have followed you to the new vendor if you had switched without finding it. Fix the mapping, keep the attributed defect log running, and revisit the vendor question in two quarters with numbers that now mean something.
