Skip to content

Glossary · Evidence & data

Precision and Recall

Precision and recall measure the quality of a classification or matching process: precision is the share of items it flagged that are correct, and recall is the share of all correct items that it flagged.

Publisher: Altss LLCContent modified
ALTSS-DATA-037

For record matching, precision asks: of the pairs we linked, how many really are the same entity? Recall asks: of the pairs that really are the same entity, how many did we link? Making a matcher stricter usually raises precision and lowers recall. In entity resolution a false positive is a false merge of two different entities, and a false negative is a duplicate left in place.

Formal definition

With true positives (TP), false positives (FP) and false negatives (FN): precision = TP / (TP + FP), recall = TP / (TP + FN), and the F-measure is their harmonic mean. Hand and Christen (2018) note that in record linkage, high class imbalance makes accuracy meaningless, which is why precision and recall are used.

Formulas

Precision

Precision = TP / (TP + FP)
TP
items flagged (for example pairs linked) that are correct
FP
items flagged that are not correct (false positives, for example false merges)

Also called positive predictive value. Undefined when nothing is flagged.

Recall

Recall = TP / (TP + FN)
FN
correct items that were not flagged (false negatives, for example missed duplicates)

Also called sensitivity or true positive rate. Requires knowing what was missed, which usually needs a labelled sample drawn from outside the flagged set.

F1 score

F1 = 2PR / (P + R)
P
precision
R
recall

Harmonic mean of precision and recall; weights them equally. Undefined when precision or recall is undefined; 0 when both are 0.

F-beta score

Fβ = (1 + β²) × P × R / (β² × P + R)
\beta
weight on recall relative to precision: beta > 1 favours recall, beta < 1 favours precision

Hand and Christen show the F-measure can be rewritten as a weighted sum of precision and recall whose weights depend on the linkage method being evaluated; they argue that the relative importance of precision and recall should be set by the problem and the user, not by the method, and suggest alternative measures. Reporting P and R separately avoids the problem.

The confusion matrix

Truly the same entityTruly different entities
Linked by the matcherTrue positive (TP)False positive (FP): a false merge
Not linkedFalse negative (FN): a missed duplicateTrue negative (TN)

Precision uses the top row; recall uses the left column. True negatives, which dominate any matching problem, appear in neither.

Why accuracy is not used

Record linkage compares far more non-matching pairs than matching ones. Hand and Christen (2018) point out that this class imbalance makes accuracy and misclassification rate meaningless for judging linked records, as the worked example shows. The same reasoning applies to other rare-event detection, such as flagging relevant filings or extracting events from news: a detector that flags nothing looks almost perfect on accuracy.

Thresholds and the trade-off

Most matchers produce a confidence score and link pairs above a threshold. Moving the threshold trades precision against recall. The Fellegi-Sunter model of probabilistic matching, covered under entity resolution, uses two thresholds: pairs above the upper one are linked, pairs below the lower one are not, and pairs between go to clerical review. The thresholds are set from the error rates the user will tolerate. Precision and recall are the measures of those error rates.

Measuring in practice

  • Ground truth. Both measures need labelled pairs. Registry identifiers, such as the legal entity identifier, can supply a truth set where they exist; elsewhere, reviewers label a sample.
  • Sample design. A sample drawn only from linked pairs estimates precision but says nothing about recall. Recall needs a sample of unlinked pairs or an independent list of known duplicates.
  • Pairs or clusters. Matching that groups records into clusters can be scored on pairs or on whole clusters; results differ and the choice should be stated.
  • Context. A published figure states the dataset, the threshold, the sample, the period and the date. Precision and recall from different datasets are not comparable.

Which error costs more in private-markets data

A false merge combines two different organisations: one allocator's commitments are attributed to another, outreach reaches the wrong institution, and figures are summed across entities that should be separate. It is hard to see and hard to undo. A missed duplicate splits one organisation across records, a failure of uniqueness, one of the data quality dimensions, which causes double counting and fragmented history but is usually visible and easy to fix. For merges, practitioners therefore favour precision; for review queues, which a person will check, recall matters more.

How Altss applies this (Altss methodology)

In Altss's entity-resolution methodology, no similarity score merges records. Two organisation records are merged only on a shared registry identifier of the same scheme, scope and jurisdiction, or a filing or registry record that names both as one entity; two person records are resolved automatically only on deterministic identity evidence or on several independent strong signals that agree, and otherwise stay separate. Precision and recall of a probabilistic matcher therefore measure how well it finds and ranks candidate pairs for registry lookup and review, not how well it merges. A match score has derivation status DERIVED and is used only to rank candidates: it is never an evidence origin or a validation status, and never evidence of identity. Any published precision or recall figure would state its population, sample size, period and date; none is implied here. See the Entity Resolution Methodology.

Worked examples

Illustrative matcher at a strict threshold

A deduplication run compares 1,000,000 candidate record pairs, of which 160 are true duplicates. At a strict threshold it links 120 pairs: 96 true (TP), 24 false (FP); it misses 64 duplicates (FN). Precision = 96 / 120 = 0.80. Recall = 96 / 160 = 0.60. F1 = 0.686.

The same matcher at a looser threshold

At a looser threshold it links 200 pairs: 140 true, 60 false, and misses 20. Precision = 140 / 200 = 0.70. Recall = 140 / 160 = 0.875. F1 = 0.778.

F1 prefers the looser threshold. F0.5, which weights precision more, prefers the strict one (0.750 against 0.729). The choice of measure is a choice about which error costs more.

Accuracy would have told nothing: the strict run is 99.991% "accurate", and a matcher that links nothing at all scores 99.984%, because almost every candidate pair is a non-match.

Examples are illustrative; figures are not market data.

Not the same as

  • Confidence Score: A confidence or match score is assigned to each item; precision and recall evaluate the decisions made from those scores across a labelled sample.
  • Data Quality Dimensions: The accuracy dimension of data quality asks whether stored values match reality; precision and recall evaluate a matcher's decisions, leaving out the true negatives that make classification accuracy uninformative.

Common mistakes

  • Reporting accuracy for a matching or detection task where non-matches dominate.
  • Reporting F1 alone, which hides whether precision or recall is weak.
  • Estimating recall from a sample of linked pairs only.
  • Comparing precision and recall across different datasets or thresholds as if they were the same test.
  • Ignoring errors in the ground truth itself.
  • Confusing precision of a classifier with precision in the sense of numerical granularity (a date known only to the year).

Edge cases

  • If nothing is flagged, precision is undefined (0/0), not zero or one.
  • Pairwise scoring can count one wrong cluster of five records as several errors; cluster-level scoring counts it once.
  • Entities that legally merge or split over time change the ground truth; labels need an as-of date.
  • A matcher can have high precision on a sample dominated by easy pairs and lower precision on hard cases such as related entities with similar names.

Questions

What is a good F1 score?

There is no universal threshold. It depends on how rare the matches are, the cost of each type of error and the difficulty of the data. Report precision and recall separately with the sample and date.

Why not use accuracy?

In matching, almost every candidate pair is a non-match, so even a matcher that links nothing scores close to 100% accuracy.

Sources

  1. A note on using the F-measure for evaluating record linkage algorithms. David Hand; Peter Christen, Statistics and Computing (Springer), Vol. 28, p. 539 onward (2018); published online 2017. Status: Published (checked 2026-10-01). Abstract — supports: Record linkage as classification of pairs; class imbalance makes accuracy meaningless; precision, recall and F-measure; F-measure as a method-dependent weighted sum; relative weight should be set by the problem and user
  2. A Theory for Record Linkage. Ivan P. Fellegi; Alan B. Sunter, Journal of the American Statistical Association, Vol. 64(328), pp. 1183-1210. Status: Published (paywalled) (checked 2026-10-01). Vol. 64(328), pp. 1183-1210 — supports: Link, possible link (clerical review) and non-link decisions with error rates fixed in advance
4 terms

Concept record

Concept ID
ALTSS-DATA-037
Classification
Evidence & data
Topics
Private markets data & OSINT
Version
2.0.0
Last reviewed
Structured data
JSON