Skip to main content
Compliance

Who really owns your supplier

Ownership questions are answered with entity resolution across imperfect registries. The model is the smallest part of that.

Practitioner Note Drawn from open engineering practice and published literature. No client data, proposal content, or program-office discussion appears here.

The question sounds simple and is not: who controls this supplier. Organizations ask it for procurement policy, for regulatory screening, for concentration risk, and for continuity planning. The answer requires stitching together registries that disagree with each other, and no amount of model quality substitutes for getting that stitching right.

Where the difficulty actually lives

Corporate records are maintained jurisdiction by jurisdiction, in different formats, with different definitions of control, updated on different schedules, and with no shared identifier. The same entity appears as three spellings, two addresses, and one transliteration that does not round-trip.

A language model reading these records will produce a fluent, confident hierarchy. Whether that hierarchy is correct depends almost entirely on whether the entity resolution beneath it was correct, and the fluency of the output carries no information about that.

The output is a clean org chart either way. Only the plumbing decides whether it is the real one.

The resolution problem, concretely

Two records name a company with a one-character difference and share a partial address. Are they one entity or two? Getting this wrong in one direction merges unrelated firms and invents a relationship. Getting it wrong in the other splits a real parent from its subsidiary and hides one.

  • Match on multiple independent signals, never name similarity alone
  • Keep the match evidence, so a disputed link can be examined rather than re-litigated
  • Represent uncertainty as a state, not by picking the more likely option silently
  • Version the graph, because ownership changes and yesterday's answer needs to stay reproducible
  • Record registry retrieval dates, so staleness is visible in the answer

Control is not the same as ownership

Percentage ownership is the easiest thing to extract and frequently the wrong question. Control arrives through board composition, golden shares, financing covenants, exclusive supply arrangements, and licensing terms — none of which appear as a percentage in a registry.

Systems that report only equity chains will therefore return clean answers that miss the arrangements the policy was written to catch. Where the underlying documents exist, extracting control indicators as separate, span-linked findings is more useful than a single ownership number.

Where ownership answers actually break

Entity resolution across inconsistent registries
90%
Control that arrives without equity
84%
Jurisdictional coverage gaps
78%
Refresh and change detection
64%
Traversal depth decisions
50%
Name matching
28%

Relative weighting from delivery practice, not a measured statistic — shown to rank where attention belongs.

A note on how these findings should read

Screening output describes corporate facts: ownership, funding, control, joint ventures, licensing. It should be written that way — attributed to the registry or statute it came from, dated, and free of inference the records do not support.

That is not only a fairness point, though it is that. It is an accuracy point. A finding phrased as a corporate fact can be checked. A finding phrased as a characterization cannot, and it will not survive the first challenge from a supplier who has the documents.

What a good system produces

Not a verdict. A dossier: the resolved entity, the evidence for each link, the retrieval dates, the confidence on each edge, the control indicators found and where, and an explicit list of jurisdictions where no record was available.

That last list is the honest part. Coverage gaps are the normal condition in this work, and a system that reports a complete hierarchy without reporting where it could not look has hidden the most important caveat in the file.

The registry landscape, honestly described

Ownership work depends on public and commercial sources that vary enormously in completeness. Understanding what each can and cannot tell you prevents a great deal of misplaced confidence.

Source typeGenerally gives youGenerally does not
Company registriesLegal existence, registered address, officers, filingsConsistent beneficial ownership; identifier interoperability
Beneficial ownership registersDeclared controlling interests where the regime requires themCoverage outside those regimes; verification of declarations
Securities filingsDetailed structure for listed entities and their subsidiariesAnything about private firms, which is most of a supply base
Commercial data providersPre-resolved hierarchies with broad coverageTheir resolution evidence — you inherit decisions you cannot inspect
Contract and diligence documentsControl terms that appear in no registry at allSystematic coverage; you only have what you collected

The commercial provider row deserves emphasis. Buying a resolved hierarchy is often the right call, and it means adopting another organization's resolution decisions without access to the evidence behind them. That is acceptable for screening and weak for a determination someone will contest, which argues for treating provider hierarchies as a strong prior rather than as an answer.

What match evidence looks like when you keep it

The recommendation to retain match evidence is easy to state and worth making concrete. For each link between two records, the system stores the signals that supported it and the signals that did not.

Consider a link asserted between a supplier record and a registry entity. The evidence might record an exact identifier match on a national registration number, a name match at high similarity after normalization, an address match on street and postal code but not on suite, a jurisdiction match, and a date-range overlap consistent with continuous operation. It might also record one contrary signal: a sector code that does not align.

Stored this way, a challenge becomes tractable. Someone disputing the link sees exactly what supported it, can point at the contrary signal, and the adjudication is recorded alongside. Stored as a merged golden record, the same challenge produces a research project and, frequently, a different answer than last time.

Refresh, change detection, and the reproducibility problem

Ownership is not static and neither are the registries describing it. A supplier screened at onboarding and never re-examined is screened against a world that no longer exists, and the gap grows silently.

  • Scheduled refresh tied to supplier criticality rather than a uniform interval
  • Event-driven refresh on contract renewal, scope expansion, or an external signal
  • Change detection that reports what moved since the last look, not just the current state
  • Versioned graph snapshots, so a determination made in March is explicable in terms of March's data
  • Retrieval dates surfaced in output, so a consumer of the answer can see how old it is

The versioning requirement is the one teams skip and later need. A decision defended two years after the fact has to be evaluated against what was knowable at the time, and a system that only holds current state cannot support that.

Control indicators worth extracting

Because control frequently arrives through arrangements rather than equity, the documents that describe those arrangements are worth systematic extraction where they exist — diligence files, joint venture agreements, financing documents, distribution and licensing agreements.

The indicators that matter are board appointment rights, veto or consent rights over defined actions, share classes with disproportionate voting, financing covenants that transfer control on breach, exclusive supply or distribution arrangements that create dependency, and licensing terms where key technology is held elsewhere.

Each extracted as a span-linked finding, attributed to the document and clause it came from. This is a good application for a language model — the language is varied, the concepts are stable, and the output is a structured finding a human confirms. It is not a good application for summarization, for the same reason summarization fails anywhere the output has to be checkable.

Writing findings that survive a challenge

Screening output frequently ends up in front of the supplier, who has the documents and a commercial interest in disputing the conclusion. That prospect should shape how findings are written.

A finding phrased as a corporate fact with a citation — the registry shows a named entity holding a stated percentage, filed on a stated date — is checkable and either right or wrong. A finding phrased as a characterization is neither checkable nor defensible, and it converts a factual conversation into an argument about judgment.

Where a rule or statute defines the standard, cite it and attribute it. Where the data is incomplete, say which jurisdictions were searched and which were not. A report that names its own coverage gaps is more credible, not less, and it prevents the far worse outcome of a complete-looking hierarchy that quietly omitted a jurisdiction where the relevant entity sits.

Tier one is where the data stops and the risk does not

Most screening programs resolve direct suppliers well and go no further, because tier-two relationships are not disclosed anywhere the buyer can reach. That boundary is a data limit rather than a risk limit, and the distinction is worth making explicit in the report.

Where second-tier visibility matters, the practical routes are contractual rather than analytical: disclosure obligations in the supply agreement, attestation at renewal, and the right to request the supplier's own screening output. None of these are perfect and all of them beat inferring a hierarchy the registries cannot support.

The reporting discipline follows. A dossier that resolves tier one and says so is honest. A dossier that resolves tier one and is silent about tier two invites the reader to assume coverage that does not exist, which is how a screening program creates false comfort rather than reducing risk.

Turning a finding into an action

A resolved ownership graph produces findings faster than most organizations can act on them, and an unactioned finding is a documented awareness of a risk that nothing was done about — a worse position than not having looked.

  • Define in advance which findings require escalation and to whom
  • Set a response clock, so an open finding has an age and a due date
  • Record the disposition — accepted, mitigated, exited — with the rationale
  • Distinguish findings that block onboarding from those that require monitoring
  • Feed dispositions back so the same relationship is not re-adjudicated every cycle

That last item is what keeps the program sustainable. Without it, every refresh regenerates the same findings, reviewers learn the queue is noise, and the genuinely new relationship gets the same attention as the one cleared three times already.

What to hand the supplier

Suppliers challenged on an ownership finding respond far better to a factual extract than to a conclusion. Providing the registry records relied on, the retrieval dates, and the specific link in question converts a confrontation into a data correction, which is usually what it actually is.

It also produces better data. Suppliers hold documents no registry contains, and a factual, non-accusatory request is the cheapest way to obtain them. Programs that lead with a determination get lawyers; programs that lead with "here is what the record shows, please correct it" get filings.

Frequently asked questions

Can this be fully automated?

The retrieval, resolution, and evidence assembly can. The adjudication of a contested link benefits from a human, and the system should be built to route those rather than resolve them quietly.

How often does the graph need refreshing?

Tied to the decision it supports. A one-time screen at onboarding with no refresh produces answers that are correct on the day and progressively wrong afterward.

1 business day response

Working on something like this?

We build systems where every figure is executed against the real record, every sentence carries the source it came from, and the system says so when the data does not support an answer.

Start a conversationCapabilitiesRead more insights →
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE