Glossary · Evidence & data
Entity Resolution
Also called: record linkage · entity matching · identity resolution · merge-purge
Entity resolution is the process of deciding which records, within or across datasets, refer to the same real-world entity (a company, fund, institution or person) and linking or merging them while keeping distinct entities apart.
One pension plan can appear as its full legal name, an acronym and the name of its investment office. Those records describe one entity and belong together. A similarly named plan for a different group of employees, run from the same building and sharing a web domain, is a different allocator and must stay separate. Entity resolution makes that call. It answers only whether two records are the same entity; how distinct entities are connected (relation) and what each one does (role) are separate questions.
Formal definition
Fellegi and Sunter (1969) formalised record linkage as deciding, for each pair of records, between link, possible link (clerical review) and non-link, using the probabilities of the observed field agreements among true matches (m) and among non-matches (u), with the error rates fixed in advance. Hand and Christen (2018) describe record linkage as classifying record pairs as matches, referring to the same real-world entity, or non-matches.
Formula
Fellegi-Sunter match weight
- mi
- probability that field i agrees when the two records are the same entity
- ui
- probability that field i agrees when the two records are different entities (chance agreement)
- W
- total weight of the pair; compared with an upper threshold (link) and a lower threshold (non-link), with pairs in between sent to clerical review
The weights are log-likelihood ratios; base 2 is common but any base works if used consistently. The model assumes fields agree independently given match status. The m and u probabilities are estimated from labelled pairs or from the data, and the thresholds are set from the error rates the user accepts.
Identity, relation and role
Three questions are kept apart:
- Identity: are these two records the same entity?
- Relation: if not, how are the two entities connected (subsidiary of, general partner of, advised by, invests for, succeeded by)? See relationship graph and legal entity hierarchy.
- Role: what does each entity do (allocates capital, manages funds, advises, operates a business)?
Answering one never answers another. Telling apart distinct entities with similar names (identity disambiguation) is the other side of the identity question. Linking a firm to its subsidiaries, funds or special purpose vehicles (SPVs) is relation-building, not resolution: they are distinct legal entities. A shared web domain, name stem, address or board member is evidence of a possible relation, not of identity. Metrics such as assets or commitments belong to the entity that reports them and are not added up across a family of related entities.
How resolution works
- Normalise for comparison: strip legal-form suffixes, punctuation and case; transliterate; expand abbreviations. The names as stated in sources are kept; normalised strings are only used to compare.
- Generate candidates (blocking): compare only plausible pairs, grouped by keys such as name tokens, jurisdiction or identifiers, so the number of comparisons stays manageable.
- Compare fields: record agreement, partial agreement or disagreement per field.
- Score: apply deterministic matching rules (exact agreement on specified fields) and, where useful, a probabilistic matching model.
- Decide: link, send to review, or keep apart.
- Record: write the decision with its evidence, the person or system that made it, and the date, so it can be audited and reversed.
Registry identifiers as deterministic keys
The strongest identity evidence is an identifier assigned by a statutory register or regulator, because the scheme's rules make it unique within its scope:
- Legal entity identifier (LEI): a 20-character code under standard 17442 of the International Organization for Standardization (ISO); each LEI identifies exactly one legal entity. Its "who owns whom" (Level 2) data describes relations between entities, not identity.
- Central Index Key (CIK) of the Securities and Exchange Commission (SEC): a unique number assigned to each filer on the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, corporations and individuals alike; CIKs are not recycled. One filer can hold more than one, and an adviser's Form ADV lists all of its CIK numbers.
- Central Registration Depository (CRD) number: assigned through the CRD system operated by the Financial Industry Regulatory Authority (FINRA), or through the Investment Adviser Registration Depository (IARD) for advisers. The CRD number on an adviser's Form ADV is the adviser's own, not that of an officer, employee or affiliate.
- Employer identification number (EIN): a nine-digit number from the Internal Revenue Service (IRS) identifying a tax account. An entity should have only one, though duplicates can be issued in error. Employee benefit plans are identified by the sponsor's EIN plus a three-digit plan number, so an EIN match alone does not make a plan its sponsor.
- French register numbers: the legal-unit number (SIREN) has nine digits and is never reused; the establishment number (SIRET) has 14 digits, made of the SIREN plus a five-digit suffix. Two SIRETs with one SIREN are two establishments of one legal unit.
- UK company number: allocated by the registrar to every company, including registered overseas companies; the registrar can change the form of numbers, with notice.
Each identifier merges only things of the kind it identifies, within its own scheme and jurisdiction. Different numbers from one scheme do not always mean different entities: an adviser can report more than one CIK and additional CRD numbers, and the IRS warns that duplicate EINs can be issued, so in those schemes different numbers alone do not prove two entities. A web domain is not an identifier: domains are shared across affiliates and change hands.
Probabilistic matching
Where no shared registry identifier exists, descriptive fields are compared. The Fellegi-Sunter model scores each candidate pair by how much more likely its pattern of agreement is among true matches than among non-matches (formula above), and sets two thresholds so that the rates of false links and missed links stay within limits chosen in advance. Pairs between the thresholds go to clerical review. Machine-learning matchers follow the same logic with learned features. The quality of any threshold is measured with precision and recall.
Duplicate Detection (deduplication)
Duplicate detection is entity resolution within one dataset: finding records that describe the same entity so they can be merged. Two distinctions matter:
- Duplicate vs sibling. A duplicate is a second record of one entity and is merged once identity is established; the surviving record is the canonical record for that entity. A sibling is a distinct entity in the same family (a parallel fund, an affiliate, the other plan run by the same board) and is related, never merged. Until a pair is classified, it is neither.
- Merge mechanics. A merge records the surviving record, the absorbed record, the evidence that justified it, who made it and when. The absorbed record's identifier keeps resolving to the survivor. Conflicting values are kept with their evidence rather than overwritten by a "survivorship" rule that silently picks one. Every merge can be undone, returning each claim to the record its evidence supports.
Aggressive deduplication improves headline counts and creates false identities; a false merge is usually harder to find and undo than a missed duplicate.
Private-markets traps
- Fund, general partner entity and management company often share a name; they are distinct entities with distinct identifiers.
- Parallel funds, feeder funds, alternative investment vehicles and continuation vehicles are distinct from the main fund; a successor fund is a new entity under the same manager.
- Asset owner and investment office. An endowment and the company that invests for it, or a plan and its outsourced chief investment officer (OCIO), are two entities related by "invests for" or "advises".
- Family-office groups contain many entities (holding companies, trusts, foundations) related by household or ownership, not one entity.
- Name changes and mergers. A renamed entity keeps its identifier; a merger creates a "succeeded by" relation between distinct entities.
- People. Identical names are common, and a name, an employer or a location alone never shows that two records are one person. An identifier that clearly refers to one individual, such as an individual's CRD number, can establish identity. A person's job is a dated relation to an organisation with its own validity period (see temporal validity).
How Altss applies this (Altss methodology)
Altss applies a stricter rule than classic record linkage, in which pairs above an upper threshold are linked automatically. Two organisation records are merged only when both carry the same registry identifier of the same scheme, within that scheme's identification scope and jurisdiction, or when a filing or registry record names both as one entity. No other evidence merges organisation records. Different LEIs mark distinct legal entities; in schemes where one entity can hold several numbers (CIK, CRD number, EIN), different numbers keep records apart until a register or a filing by the entity links them. Similarity of any kind (name, domain, address, phone, shared people) creates a relation candidate that is reviewed and is not published, counted or used to combine data until evidence supports a typed, dated relation.
Natural persons follow a conservative rule of their own. Two person records are resolved to one person automatically only on one of two bases: deterministic identity evidence, meaning a unique identifier that clearly refers to the same individual in both records (for example an individual's CRD number, or a unique canonical identifier issued by the source for that individual); or several independent strong signals that agree, meaning compatible full names or name variants together with further independent evidence, such as the same organisation, a compatible role, a consistent employment history, consistent geography, the same unique professional-profile web address (URL), the same verified professional contact identifier, or other independent corroborating sources, drawn from evidence chains that do not share an originating source. No single weak attribute is sufficient: a name alone, an employer alone, a location alone or an assertion by a language model alone never establishes that two records are one person. Hard evidence conflicts block automatic resolution: incompatible unique identifiers, clearly distinct professional profiles, independent evidence of two separate individuals, or identity histories that cannot reasonably belong to one person. Where the evidence is insufficient or conflicting, the records stay separate, and records are never merged to make a profile look more complete. Ambiguous cases with a high impact, such as a senior decision-maker at an allocator, go to research review before any merge, and the review is recorded under the evidence model's validation rules. The methodology publishes no numeric matching thresholds: the rule states which evidence is admissible, not a calibrated score.
Fellegi-Sunter scoring is used to rank candidate pairs for registry lookup and review; in this methodology no score band merges. Merges are recorded with the identifier, filing or evidence that justified them and are reversible. Roles are never inferred from names, and no metric is summed across an entity family. Identifier evidence is classified on the three evidence dimensions: an identifier retrieved from the issuing register has evidence origin REGULATORY_PUBLIC_RECORD and derivation status OBSERVED; the same number quoted elsewhere (DISCLOSED, OSINT_SOURCED or LICENSED_THIRD_PARTY) merges nothing until the register confirms it; a match score has derivation status DERIVED and is used only to rank candidates; and validation status is recorded per claim with its date. See the Entity Resolution Methodology.
Worked example
Illustrative candidate pair scored on three fields
Two records are compared on three fields, with illustrative probabilities:
| Field | m | u | Weight if agree | Weight if disagree |
|---|---|---|---|---|
| Normalised name | 0.95 | 0.002 | +8.89 | −4.32 |
| City | 0.90 | 0.05 | +4.17 | −3.25 |
| Founding year | 0.80 | 0.03 | +4.74 | −2.28 |
- All three agree: W = 17.80.
- Name and city agree, founding year differs: W = 10.78.
- Only the name agrees: W = 3.37.
A classic Fellegi-Sunter setup would link the first pair automatically if 17.80 cleared the upper threshold. A registry lookup can still show that the two records carry different legal entity identifiers: for example a management company and a same-named affiliate at the same address. High agreement on descriptive fields measures similarity, not identity.
Examples are illustrative; figures are not market data.
Not the same as
- Legal Entity Hierarchy: A hierarchy records ownership and control relations between distinct legal entities; resolution never merges a parent with its subsidiaries.
- Relationship Graph: A relationship graph connects distinct entities by typed relations; entity resolution determines what the nodes are.
Common mistakes
- Treating a high similarity score as proof of identity without checking the registry identifiers that are available.
- Using a web domain, a phone number or an address as if it were a unique identifier.
- Merging a fund with its manager or general partner because they share a name.
- Summing assets or commitments across related entities and presenting the total as one entity's figure.
- Merging records irreversibly or without recording the justification.
- Letting one value silently overwrite another on merge.
- Inferring what an organisation does from its name.
- Treating resolution as a one-time cleanup rather than a decision repeated as new records arrive.
Edge cases
- Two records with different numbers from a scheme in which one entity can hold several (CIK, CRD number, EIN) are kept apart but are not thereby proved to be two entities; a filing in which the filer lists both numbers as its own resolves them to one.
- A redomiciled entity can carry identifiers from two schemes. Treating them as one entity needs evidence that connects the two registrations, such as a filing that names both; similar names and addresses across the schemes do not establish it.
- A UK registered number can change if the registrar adopts a new numbering form; the old and new numbers refer to the same company.
- Identifiers can be wrong in a filing (a typo in a CRD or CIK); a single filing's identifier that conflicts with the register is checked against the register before any merge.
- An EIN on a benefit plan's filing belongs to the sponsor; the plan is identified by the EIN plus its plan number.
Questions
What is the difference between entity resolution and deduplication?
Deduplication is entity resolution within one dataset. Entity resolution also links records across datasets and must decide when similar records are different entities.
Is a matching domain name enough to merge two companies?
No. Domains are shared by affiliates and change hands. A shared domain suggests a possible relation to investigate; identity rests on a registry identifier or a filing naming both as one entity.
External standards
| Standard | Relation | Note |
|---|---|---|
| ISO 17442 (LEI) (One LEI identifies exactly one legal entity) | related |
Sources
- A Theory for Record Linkage. Ivan P. Fellegi; Alan B. Sunter, Journal of the American Statistical Association, Vol. 64(328), pp. 1183-1210. Status: Published (paywalled) (checked 2026-10-01). Vol. 64(328), pp. 1183-1210 — supports: m- and u-probabilities, agreement weights, link / possible link / non-link with upper and lower thresholds set from error rates fixed in advance (classic automatic linkage above the upper threshold)
- A note on using the F-measure for evaluating record linkage algorithms. David Hand; Peter Christen, Statistics and Computing (Springer), Vol. 28, p. 539 onward (2018); published online 2017. Status: Published (checked 2026-10-01). Abstract — supports: Record linkage as classification of pairs into matches and non-matches; evaluation by precision and recall
- Introducing the Legal Entity Identifier (LEI). GLEIF, Global Legal Entity Identifier Foundation, ISO 17442 identifier; web page accessed 2026-10-01. Status: Current (checked 2026-10-01). Introducing the LEI page — supports: LEI is a 20-character ISO 17442 code; each LEI identifies exactly one legal entity; Level 2 data is who owns whom
- EDGAR Glossary. U.S. Securities and Exchange Commission, Page last reviewed or updated 2025-01-16. Status: current (checked 2026-10-01). Central Index Key (CIK) — supports: CIK is a unique number assigned to each EDGAR filer
- Accessing EDGAR Data. U.S. Securities and Exchange Commission, Page last reviewed or updated 2024-06-26. Status: current (checked 2026-10-01). CIK paragraphs (unique to the filer, not recycled) — supports: CIKs remain unique to the filer and are not recycled
- CIK Lookup. U.S. Securities and Exchange Commission, Page last reviewed or updated 2024-11-01. Status: current (checked 2026-10-01). Page introduction (CIK definition) — supports: CIKs identify corporations and individual people who have filed with the SEC
- Central Registration Depository (CRD). FINRA, Web page accessed 2026-10-01 (no date shown). Status: current (checked 2026-10-01). Page introduction (scope of the CRD program) — supports: FINRA operates the Central Registration Depository for firm and individual registration records
- Form ADV Part 1A (paper version) - Uniform Application for Investment Adviser Registration and Report by Exempt Reporting Advisers. U.S. Securities and Exchange Commission, SEC 1707 (07-24). Status: in force (checked 2026-10-01). Item 1.D(3), Item 1.E(1)-(2); Schedule D Sec. 7.A — supports: Form ADV lists all of the adviser's CIK numbers, its own CRD number (not an officer's or affiliate's) and any additional CRD numbers; related persons are listed with their own identifiers
- Publication 1635, Employer Identification Number: Understanding Your EIN. Internal Revenue Service, Rev. 2-2014 (the revision hosted at irs.gov/pub/irs-pdf and linked from the IRS EIN page, reviewed 2026-07-17). Status: current (checked 2026-10-01). p. 2 (What is an EIN?); p. 15 — supports: EIN is a nine-digit IRS number; an entity should have only one; duplicates can arise
- Instructions for Form 5500, Annual Return/Report of Employee Benefit Plan (2025). U.S. Department of Labor (EBSA), Internal Revenue Service and Pension Benefit Guaranty Corporation, 2025 Form 5500 instructions (EFAST2). Status: current (checked 2026-10-01). Part II, line 1b — supports: A plan is identified by the sponsor EIN plus a three-digit plan number
- Siren number (definition). INSEE (Institut national de la statistique et des études économiques), Last update 2025-10-21. Status: current (checked 2026-10-01). Definition — supports: SIREN is a nine-digit identifier of a legal unit, allocated once and never reused
- Numéro Siret (définition). INSEE (Institut national de la statistique et des études économiques), Dernière mise à jour 2019-12-04. Status: current (checked 2026-10-01). Definition — supports: SIRET is a 14-digit establishment identifier made of the SIREN and a NIC
- Companies Act 2006, section 1066 - Company's registered numbers. UK Parliament (legislation.gov.uk, The National Archives), Latest available (revised), up to date with changes known to be in force on or before 2026-10-01. Status: in force (checked 2026-10-01). Sec. 1066(1)-(4), (6) — supports: Registrar allocates a registered number to every company, including registered overseas companies; the form of numbers can be changed with notice
Related terms
6 termsReferenced by
2 termsConcept record
- Concept ID
- ALTSS-DATA-024
- Classification
- Evidence & data
- Topics
- Private markets data & OSINT
- Version
- 2.0.0
- Last reviewed
- Structured data
- JSON