Skip to main content
Commercial

Replacing a vendor black box with a model you own

A vendor score you cannot inspect, retrain or move is a dependency whose price is set by your switching cost. This is how to decide whether to own it, what building really takes, and the reversible stages that let you run both until the new one has proven itself.

The renewal quote arrives and the increase is larger than last year's, which was larger than the one before. Somebody asks the obvious question: what exactly are we buying. The answer is a score. You cannot see how it is computed, you cannot retrain it on your own population, you cannot take it with you, and when a regulator or a large customer asks you to explain a decision it drove, you forward a vendor document and hope it satisfies them. That is not a purchase. It is a dependency, and its price is set by how hard it would be to leave.

This is written for the executive holding that renewal: chief data officer, chief risk officer, head of analytics. The question is not whether vendor models are bad. Many are good, and building everything yourself is a mistake. The question is which ones you should own, and that has a defensible answer you can reach with a few weeks of work.

What the dependency actually costs

The invoice is the visible cost and usually the smaller one.

The escalation is structural. A vendor whose product is embedded in your decisioning knows the cost of removing it. Renewal pricing follows switching cost, not value delivered, and switching cost grows every year the integration deepens. Nothing about that is unusual or improper; it is simply how the position works, and it is the reason to think about it before the next renewal rather than during it.

You cannot fit it to your population. A vendor model is trained on a reference population that is not yours. If your customers, exposures or geographies differ from that reference, the score is systematically miscalibrated on you, and you cannot fix it. The usual workaround is a layer of internal adjustments on top of the vendor output, which means you are already maintaining a model and getting none of the benefits of owning one.

The explanation is not yours to give. Where the law requires reasons for a decision, or a large customer's risk team asks how the score was derived, you are relaying an account you did not produce and cannot verify. Supervisors have grown less patient with this, and the burden sits with the firm using the model, not with the vendor supplying it.

The roadmap is theirs. Model updates arrive on the vendor's schedule and change your decisions with them. You may get notice and a comparison. You rarely get a veto, and you almost never get the option to stay on the old version indefinitely while you assess the new one.

The data flows outward. Scoring usually means sending your data to the vendor. What they retain, and whether it improves a model sold to your competitors, is a contract question most firms have never read closely.

Signals that a vendor model is worth replacing with one you own

You already maintain adjustment layers over the vendor output
93%
You hold outcome data the vendor never sees
90%
You must explain decisions the score drives
87%
Renewal pricing has outgrown any change in value
82%
The score is a differentiator in what you sell
78%
The vendor's proprietary data is the whole of the value
19%

Editorial weighting, illustrative rather than measured. The last row is deliberately low: where the value is data you cannot obtain, keep buying it.

When to build, and when not to

Build when the value is in the modeling and you hold the data. If the vendor's advantage is technique applied to inputs you already have, plus outcome history you own and they do not, you are in the strong position and you should be uncomfortable renting it.

Build when the decision is regulated or explained. Owning the model means owning the documentation, the validation, the reason codes and the ability to answer a supervisor without an intermediary.

Build when the score differentiates your product. If your customers buy your judgment, renting that judgment from a supplier who sells it to your competitors is a strategic position worth examining rather than a procurement decision.

Build when your population is distinctive. The more your book differs from the vendor's reference population, the more a model fitted to your own data will beat theirs, and the more the vendor's calibration works against you.

Do not build when the value is proprietary data you cannot get. If the vendor's edge is a panel, a bureau file, a consortium of contributed data or a licensed source, no amount of modeling replaces it. In that case the right move is often narrower: keep buying the data and build the model on top of it yourself, which is a different negotiation and frequently a cheaper one.

Do not build when the decision is peripheral. Building carries permanent operating and governance cost. Spend that on the models that matter to your economics and keep buying elsewhere.

Do not build when you lack outcome data. Without labels the exercise cannot start, and the honest first project is instrumenting the outcome capture, then revisiting the question in a year with data you did not have.

If the vendor's advantage is technique applied to inputs you already have, plus outcome history you own and they do not, you are in the strong position and you should be uncomfortable renting it.

What it takes, stated honestly

Four things, and the modeling is the smallest.

Data, including labels with the right timing. You need the inputs available at decision time, not the enriched record assembled afterward, and you need outcomes with the dates they became known. Every feature has to be reconstructible as of the decision moment, which is where most build projects lose their first month. Also required: enough history to cover a period of stress, because a model fitted only on benign conditions will be confidently wrong in the next bad one.

There is a subtle trap here. Your historical outcomes were observed under decisions the vendor score drove. Applications the score declined have no outcome, so your data is censored precisely where the model was most decisive. Handling this properly, by using whatever populations were approved outside the score, exception cases, or a deliberate small share of decisions made another way, is the difference between a model that works and one that looks excellent in development and fails in production.

Validation independent of the builders. Regulated firms know this discipline: a second team reviews the conceptual soundness, the data, the testing and the limitations, and its report goes to a committee. Firms without that structure should adopt it anyway, because it is what makes the model defensible and because it catches the errors the builders cannot see.

Governance and documentation. The inventory entry, the development documentation, the testing evidence, the monitoring plan, the change process, the named owner. Buying a vendor model outsources some of this. Owning it means doing it, and it is real ongoing work.

Operations. Scoring at your volume and latency, feature computation that matches training exactly, fallback behavior when an input is missing, monitoring with thresholds, and a retraining process. The gap between training and serving is where quiet failures live, and the fix is one code path used by both rather than two implementations that are supposed to agree.

Running them side by side, which is how you de-risk it

Nobody should switch a decisioning model on a date. The transition has stages and each one is reversible.

StageWhat runsWhat you learnExit if it fails
RetrospectiveNew model scored on history; vendor score is the incumbentWhether it beats the vendor on your own outcomes, by segmentStop. Nothing has changed in production
ShadowBoth score live traffic; only the vendor decidesAgreement rate, where they disagree, operational behaviorTurn off the shadow. No decision was affected
Champion and challengerNew model decides a small share, randomly assignedReal outcomes on decisions your model madeReturn the share to zero immediately
MajorityNew model decides most volume; vendor retained on a sliceWhether the advantage holds at scale and over timeShift volume back; the vendor path is still live
Own itYour model decides; the vendor is a benchmark or goneCost and control, with monitoring as the safety netContractual re-entry terms, negotiated while the relationship was still healthy

Two points on this sequence. The shadow stage is where most of the learning happens and it is cheap, so run it longer than feels necessary; the disagreement analysis, meaning the cases where the two models differ most and what happened to them, is the single most informative artifact of the whole program.

And the champion and challenger stage needs the outcome horizon to be respected. If defaults or claims emerge over many months, then a few weeks of live traffic tells you about operational behavior and nothing about accuracy. Plan the calendar around the label lag rather than around the renewal date, and if the renewal forces the issue, negotiate a shorter bridge term rather than compressing the evidence.

The exit terms nobody negotiates until it is too late

The time to negotiate leaving is while you are still a customer in good standing with an alternative under construction. A few terms are worth more than a price concession.

Scores you have already received are yours to keep and use, including for building a replacement, without a term that says otherwise. Read the current agreement for a clause preventing use of the output to develop a competing model, because these exist and they are exactly the constraint that matters here. Historical scores are returned in bulk in a usable format on termination, and the format is specified now rather than described as reasonable. Data you supplied is deleted on request with confirmation. The transition period is defined, with the option to run in parallel while you migrate, priced in advance rather than at the moment you have no alternative. And where the model changes materially, you get notice, a comparison and a defined period on the prior version.

None of these are unusual asks. All of them are much harder to get once the vendor knows you are leaving.

What ownership changes

The cost picture inverts rather than simply shrinking. The vendor arrangement is mostly recurring fees that grow. Ownership is mostly a build cost followed by lower recurring cost: infrastructure, a fraction of engineering and data science attention for monitoring and periodic retraining, and the governance work you would be doing regardless. Whether the arithmetic works depends on your volume and on how much of the value was the vendor's data rather than their modeling, and it is worth computing properly with your own numbers rather than asserting either way.

Control changes more than cost. You decide when the model changes and you can hold it stable through a period when stability matters. You can fit it to your population and recalibrate as the population moves. You can add the features you hold and they cannot see, which is often where a real accuracy advantage appears. And when someone asks why a decision was made, the answer comes from your own documentation.

The regulatory posture improves for the same reason. Supervisors expect a firm to understand the models driving its decisions, and understanding is easier to demonstrate for a model you built, validated and documented than for one you licensed. Ownership does not remove any obligation; it makes meeting them a matter of producing what you already have.

There are real losses to weigh. You take on model risk that previously sat, at least in perception, with a supplier. You lose whatever benchmarking value came from using an instrument your counterparties also use. And you acquire a permanent operating responsibility. Those are reasons to be selective about which models you own, not reasons to own none.

The architecture, so the replacement is not a second lock-in

A firm that replaces a vendor score with an in-house model wired directly into its decisioning has improved its position and repeated the original mistake in a new costume. The design goal is not a model. It is a decisioning layer where the model is a component that can be swapped.

The shape is a scoring service behind a stable interface. Callers send a decision request and receive a score, a set of reason codes, a model version identifier and a confidence indicator. Behind that interface, any number of models can run: the vendor score, your own, a simple rule as a fallback, and whatever comes next. Routing between them is configuration, not a code release, which is exactly what makes the staged migration above practical rather than theoretical.

Underneath sits a feature layer that computes inputs once and serves them to both training and production scoring through the same code path. This is the detail that prevents the most common quiet failure in deployed models, where a feature is computed one way in a training notebook and another way in the serving system, and the model degrades for reasons nobody can find because the two implementations were never compared. One definition, one implementation, used by both.

Every request and every response is logged with the model version, the feature values used, the score, the reason codes and, eventually, the outcome when it arrives. That log is four things at once: your monitoring input, your evidence for an examiner, your training data for the next version, and the record that lets you answer a customer's question about a decision made eighteen months ago. Building it is cheap at the start and expensive to reconstruct later.

The decision policy itself lives outside the model. Thresholds, overrides, segment-specific rules and the treatment of edge cases are configuration with their own version history and their own approvals, because they change far more often than the model does and they should not require a model release. Firms that bury policy inside model code end up retraining a model to change a cutoff, which is both slow and impossible to explain.

Where a replacement program actually spends its effort

Reconstructing decision-time data and handling censoring
95%
Serving infrastructure with one shared feature code path
88%
Shadow running and the disagreement analysis
85%
Validation, documentation and committee review
81%
Monitoring, retraining process and named ownership
77%
Fitting the model itself
26%

Editorial weighting, illustrative rather than measured. The last row is deliberately low: the modeling is rarely what makes or breaks a replacement.

Two conversations to have internally first

Before any of this is scoped, two internal questions decide whether it succeeds, and neither is technical.

The first is who owns the model afterward. A model without a named owner accountable for its performance, its documentation and its retraining will drift and eventually be replaced by another vendor product, at a worse price, in three years. The owner does not have to be a large team. It has to be a person with the responsibility written into their role, supported by engineering that keeps the pipeline running. Firms that skip this find that ownership defaults to whoever built it, and that person changes jobs.

The second is what happens to the relationship with the vendor. Many firms are not leaving entirely; they are moving from buying a score to buying data, or reducing to a benchmark subscription, or keeping the vendor for a segment where their reference population genuinely is better. Deciding that in advance changes how the negotiation is conducted and usually produces a better outcome than a binary renew-or-leave framing, because the vendor also prefers a smaller continuing relationship to a lost account.

How we work inside your organization

Precision Federal builds and deploys these systems. Our engineers do the data reconstruction, the feature work, the modeling, the validation support, the serving infrastructure and the monitoring, and it all runs in your environment under your control.

The first engagement is the retrospective, because it settles the question cheaply. In the first two weeks we reconstruct the decision-time data, including handling the censoring in your historical outcomes, and build the comparison against the vendor score on your own book. By week four you have a segment-level answer: where a model fitted to your data beats the incumbent, where it does not, and what the operating requirements would be. That is a decision-grade result and it is worth having whether you proceed or renew, because it also tells you what you are actually buying.

You keep everything. The model, the code, the features, the documentation, all yours in your repositories under a written assignment. It runs in your cloud account under your identity provider, and your data stays inside your boundary. There is no platform license from us and nothing that stops working when the engagement ends, which would be a strange thing to build for a client who is replacing a dependency. We document for the people who will own it and we work beside your team through handover.

Pricing is fixed-price milestones where scope is clear, each tied to an acceptance test agreed in advance, so you know the cost before work starts and pay for a result. Where the work is genuinely exploratory we use a committed team at a fixed monthly rate.

We also build and deploy inside U.S. federal agencies, where a system earns an authorization to operate, produces evidence continuously, handles controlled unclassified information and meets accessibility conformance. That is why we treat validation support, documentation and monitoring as part of the build rather than as a phase that follows it. A model risk committee and an authorizing official ask the same question in different words.

The first step is one email with a one-page brief: what the vendor model decides, what outcome data you hold, and when the renewal lands. We return a scoped, priced statement of work.

Bottom line

Own the model when the value is modeling on data you already hold, when you must explain the decisions, when the score is part of what you sell, or when your population differs from the vendor's reference. Keep buying when the value is data you cannot otherwise obtain, when the decision is peripheral, or when you have no outcome history yet. The way to find out which case you are in is a retrospective comparison on your own book, which costs weeks rather than quarters. Then move in reversible stages: retrospective, shadow, a small live share, majority, ownership. And negotiate the exit terms now, while you are still a customer the vendor wants to keep.

Frequently asked questions

When should a company build its own model instead of licensing a vendor score?

When the vendor's advantage is modeling technique applied to inputs you already hold, and you own outcome history they never see. When the decision is regulated or must be explained to customers, since owning the model means owning the documentation and the reason codes. When the score differentiates what you sell, and renting it from a supplier who also sells to competitors is a strategic question. And when your population differs enough from the vendor's reference that their calibration works against you.

When is it a mistake to replace a vendor model?

When the real value is proprietary data you cannot obtain, such as a bureau file, a contributed consortium or a licensed panel. In that case the better move is often to keep buying the data and build the model on top of it yourself. Also when the decision is peripheral to your economics, since ownership carries permanent operating and governance cost that belongs on the models that matter. And when you lack outcome data, where the honest first project is instrumenting outcome capture and revisiting the question later.

How do you compare a new model against a vendor score fairly?

Retrospectively on your own outcomes first, using data reconstructed as of the decision moment rather than the enriched record assembled afterward. Report by segment, not only in aggregate, because the interesting result is usually that you beat the vendor in some segments and not others. Account for censoring: applications the vendor score declined have no outcome, so use approved populations, exception cases, or decisions made outside the score. Then confirm live in shadow before any decision depends on it.

How do you migrate off a scoring vendor without risk?

In reversible stages. Retrospective comparison on history, where nothing in production changes. Shadow mode, where both score live traffic and only the incumbent decides, which is where most of the learning happens. A small randomly assigned share of live decisions, planned around the horizon at which outcomes actually emerge rather than around a renewal date. Then majority volume with the vendor retained on a slice. Full ownership last, with monitoring as the safety net and negotiated re-entry terms behind you.

What contract terms matter most with a model vendor?

The right to keep and use scores you have already received, including for developing a replacement, and the absence of a clause forbidding exactly that. Bulk return of your historical scores on termination in a format specified now. Deletion of data you supplied, with confirmation. A defined transition period with parallel running priced in advance. And notice, comparison and a stated period on the prior version whenever the model changes materially. Negotiate all of these while the relationship is healthy.

1 business day response

Considering replacing a vendor scoring model?

We build the retrospective comparison, the model, the serving layer and the monitoring, all in your cloud and yours to keep. Send a one-page brief and we return a scoped, priced statement of work.

How we workMore insights →Email an engineer or email bo@precisionfederal.com
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE