The question sounds simple and is not: who controls this supplier. Organizations ask it for procurement policy, for regulatory screening, for concentration risk, and for continuity planning. The answer requires stitching together registries that disagree with each other, and no amount of model quality substitutes for getting that stitching right.
Where the difficulty actually lives
Corporate records are maintained jurisdiction by jurisdiction, in different formats, with different definitions of control, updated on different schedules, and with no shared identifier. The same entity appears as three spellings, two addresses, and one transliteration that does not round-trip.
A language model reading these records will produce a fluent, confident hierarchy. Whether that hierarchy is correct depends almost entirely on whether the entity resolution beneath it was correct, and the fluency of the output carries no information about that.
The output is a clean org chart either way. Only the plumbing decides whether it is the real one.
The resolution problem, concretely
Two records name a company with a one-character difference and share a partial address. Are they one entity or two? Getting this wrong in one direction merges unrelated firms and invents a relationship. Getting it wrong in the other splits a real parent from its subsidiary and hides one.
- Match on multiple independent signals, never name similarity alone
- Keep the match evidence, so a disputed link can be examined rather than re-litigated
- Represent uncertainty as a state, not by picking the more likely option silently
- Version the graph, because ownership changes and yesterday's answer needs to stay reproducible
- Record registry retrieval dates, so staleness is visible in the answer
Control is not the same as ownership
Percentage ownership is the easiest thing to extract and frequently the wrong question. Control arrives through board composition, golden shares, financing covenants, exclusive supply arrangements, and licensing terms — none of which appear as a percentage in a registry.
Systems that report only equity chains will therefore return clean answers that miss the arrangements the policy was written to catch. Where the underlying documents exist, extracting control indicators as separate, span-linked findings is more useful than a single ownership number.
Where ownership answers actually break
Relative weighting from delivery practice, not a measured statistic — shown to rank where attention belongs.
A note on how these findings should read
Screening output describes corporate facts: ownership, funding, control, joint ventures, licensing. It should be written that way — attributed to the registry or statute it came from, dated, and free of inference the records do not support.
That is not only a fairness point, though it is that. It is an accuracy point. A finding phrased as a corporate fact can be checked. A finding phrased as a characterization cannot, and it will not survive the first challenge from a supplier who has the documents.
What a good system produces
Not a verdict. A dossier: the resolved entity, the evidence for each link, the retrieval dates, the confidence on each edge, the control indicators found and where, and an explicit list of jurisdictions where no record was available.
That last list is the honest part. Coverage gaps are the normal condition in this work, and a system that reports a complete hierarchy without reporting where it could not look has hidden the most important caveat in the file.
The registry landscape, honestly described
Ownership work depends on public and commercial sources that vary enormously in completeness. Understanding what each can and cannot tell you prevents a great deal of misplaced confidence.
| Source type | Generally gives you | Generally does not |
|---|---|---|
| Company registries | Legal existence, registered address, officers, filings | Consistent beneficial ownership; identifier interoperability |
| Beneficial ownership registers | Declared controlling interests where the regime requires them | Coverage outside those regimes; verification of declarations |
| Securities filings | Detailed structure for listed entities and their subsidiaries | Anything about private firms, which is most of a supply base |
| Commercial data providers | Pre-resolved hierarchies with broad coverage | Their resolution evidence — you inherit decisions you cannot inspect |
| Contract and diligence documents | Control terms that appear in no registry at all | Systematic coverage; you only have what you collected |
The commercial provider row deserves emphasis. Buying a resolved hierarchy is often the right call, and it means adopting another organization's resolution decisions without access to the evidence behind them. That is acceptable for screening and weak for a determination someone will contest, which argues for treating provider hierarchies as a strong prior rather than as an answer.
What match evidence looks like when you keep it
The recommendation to retain match evidence is easy to state and worth making concrete. For each link between two records, the system stores the signals that supported it and the signals that did not.
Consider a link asserted between a supplier record and a registry entity. The evidence might record an exact identifier match on a national registration number, a name match at high similarity after normalization, an address match on street and postal code but not on suite, a jurisdiction match, and a date-range overlap consistent with continuous operation. It might also record one contrary signal: a sector code that does not align.
Stored this way, a challenge becomes tractable. Someone disputing the link sees exactly what supported it, can point at the contrary signal, and the adjudication is recorded alongside. Stored as a merged golden record, the same challenge produces a research project and, frequently, a different answer than last time.
Refresh, change detection, and the reproducibility problem
Ownership is not static and neither are the registries describing it. A supplier screened at onboarding and never re-examined is screened against a world that no longer exists, and the gap grows silently.
- Scheduled refresh tied to supplier criticality rather than a uniform interval
- Event-driven refresh on contract renewal, scope expansion, or an external signal
- Change detection that reports what moved since the last look, not just the current state
- Versioned graph snapshots, so a determination made in March is explicable in terms of March's data
- Retrieval dates surfaced in output, so a consumer of the answer can see how old it is
The versioning requirement is the one teams skip and later need. A decision defended two years after the fact has to be evaluated against what was knowable at the time, and a system that only holds current state cannot support that.
Control indicators worth extracting
Because control frequently arrives through arrangements rather than equity, the documents that describe those arrangements are worth systematic extraction where they exist — diligence files, joint venture agreements, financing documents, distribution and licensing agreements.
The indicators that matter are board appointment rights, veto or consent rights over defined actions, share classes with disproportionate voting, financing covenants that transfer control on breach, exclusive supply or distribution arrangements that create dependency, and licensing terms where key technology is held elsewhere.
Each extracted as a span-linked finding, attributed to the document and clause it came from. This is a good application for a language model — the language is varied, the concepts are stable, and the output is a structured finding a human confirms. It is not a good application for summarization, for the same reason summarization fails anywhere the output has to be checkable.
Writing findings that survive a challenge
Screening output frequently ends up in front of the supplier, who has the documents and a commercial interest in disputing the conclusion. That prospect should shape how findings are written.
A finding phrased as a corporate fact with a citation — the registry shows a named entity holding a stated percentage, filed on a stated date — is checkable and either right or wrong. A finding phrased as a characterization is neither checkable nor defensible, and it converts a factual conversation into an argument about judgment.
Where a rule or statute defines the standard, cite it and attribute it. Where the data is incomplete, say which jurisdictions were searched and which were not. A report that names its own coverage gaps is more credible, not less, and it prevents the far worse outcome of a complete-looking hierarchy that quietly omitted a jurisdiction where the relevant entity sits.
Tier one is where the data stops and the risk does not
Most screening programs resolve direct suppliers well and go no further, because tier-two relationships are not disclosed anywhere the buyer can reach. That boundary is a data limit rather than a risk limit, and the distinction is worth making explicit in the report.
Where second-tier visibility matters, the practical routes are contractual rather than analytical: disclosure obligations in the supply agreement, attestation at renewal, and the right to request the supplier's own screening output. None of these are perfect and all of them beat inferring a hierarchy the registries cannot support.
The reporting discipline follows. A dossier that resolves tier one and says so is honest. A dossier that resolves tier one and is silent about tier two invites the reader to assume coverage that does not exist, which is how a screening program creates false comfort rather than reducing risk.
Turning a finding into an action
A resolved ownership graph produces findings faster than most organizations can act on them, and an unactioned finding is a documented awareness of a risk that nothing was done about — a worse position than not having looked.
- Define in advance which findings require escalation and to whom
- Set a response clock, so an open finding has an age and a due date
- Record the disposition — accepted, mitigated, exited — with the rationale
- Distinguish findings that block onboarding from those that require monitoring
- Feed dispositions back so the same relationship is not re-adjudicated every cycle
That last item is what keeps the program sustainable. Without it, every refresh regenerates the same findings, reviewers learn the queue is noise, and the genuinely new relationship gets the same attention as the one cleared three times already.
What to hand the supplier
Suppliers challenged on an ownership finding respond far better to a factual extract than to a conclusion. Providing the registry records relied on, the retrieval dates, and the specific link in question converts a confrontation into a data correction, which is usually what it actually is.
It also produces better data. Suppliers hold documents no registry contains, and a factual, non-accusatory request is the cheapest way to obtain them. Programs that lead with a determination get lawyers; programs that lead with "here is what the record shows, please correct it" get filings.
Frequently asked questions
The retrieval, resolution, and evidence assembly can. The adjudication of a contested link benefits from a human, and the system should be built to route those rather than resolve them quietly.
Tied to the decision it supports. A one-time screen at onboarding with no refresh produces answers that are correct on the day and progressively wrong afterward.
