---
title: "Confidence Score | Altss Glossary"
description: "A confidence score is a stated measure, numeric or ordinal, of how strongly the available evidence supports a claim, a match or a judgment; its meaning is…"
canonical: "https://altss.com/glossary/confidence-score"
---

Glossary · Evidence & data

# Confidence Score

Also called: confidence level · confidence rating

A confidence score is a stated measure, numeric or ordinal, of how strongly the available evidence supports a claim, a match or a judgment; its meaning is set by the method that produces it.

Publisher: Altss LLCContent modified 2026-10-01

ALTSS-DATA-020

Different systems use the phrase for different things: a model's estimate that two records describe the same company, an analyst's confidence in the basis for a judgment, or a data provider's label for how well a field is supported. A score is not a probability unless it has been calibrated against checked outcomes, it is not comparable across providers that use different methods, and it does not say whether a fact is still current.

## Three things a confidence score can mean

- **Likelihood**: the probability that an event happens or that a claim is true ("roughly even chance", "0.7").

- **Confidence in the basis**: how strong the evidence and reasoning behind a judgment are, whatever the judgment says.

- **Model score**: the output of a classifier or matching model, such as a probabilistic match weight for a candidate pair of records in [entity resolution](https://altss.com/glossary/entity-resolution).

US intelligence analytic standards (ICD 203) keep the first two apart. Analytic products should indicate both the likelihood of an event and the analyst's confidence in the basis for that judgment; confidence may rest on the logic and evidentiary base, including the quantity and quality of source material. A confidence level and a degree of likelihood must not be combined in the same sentence.

## Standard likelihood language

ICD 203 fixes the words analysts may use for likelihood and the probability range each word carries:

| Term | Alternative term | Probability |
| --- | --- | --- |

| almost no chance | remote | 1–5% |

| very unlikely | highly improbable | 5–20% |

| unlikely | improbable | 20–45% |

| roughly even chance | roughly even odds | 45–55% |

| likely | probable | 55–80% |

| very likely | highly probable | 80–95% |

| almost certain(ly) | nearly certain | 95–99% |

The table governs likelihood, not confidence. Its value is that a reader can translate words into numbers and check them later. Irwin and Mandel (2019) make a related argument about grading intelligence information: they propose numeric probability estimates of how likely the information is to be accurate, re-evaluated as new information arrives, in place of rigid all-purpose schemes.

## Ordinal labels, scores and probabilities

An **ordinal label** (high, moderate, low) orders claims by strength of evidence. The distance between labels is undefined, so labels cannot be averaged or summed. A **numeric score** from a model is often not a probability: a matching model that outputs 0.9 for a pair has not shown that 90% of such pairs are true matches. A score becomes a statement about observed accuracy only through **calibration**: checking a sample of scored items against fresh evidence and comparing the confirmation rate with the stated level, for a stated population and period. For matching models, the effect of a score threshold is measured with [precision and recall](https://altss.com/glossary/precision-and-recall).

## Reading a confidence score

Before relying on a score, ask:

- **What does it attach to?** A claim, a whole record, a match between records, or a source. A record-level score hides weak fields behind strong ones.

- **What goes into it?** [Source reliability](https://altss.com/glossary/source-reliability), corroboration, how the value was produced (observed, derived, estimated), the age of the evidence.

- **When was it assessed?** A score without a date cannot be checked against later evidence.

- **Has it been calibrated, on what population, and when?** Without that, it orders claims but does not measure accuracy.

- **Does it say anything about currency?** Usually not. A well-supported statement about 2023 is still a statement about 2023. See [data freshness](https://altss.com/glossary/data-freshness).

## How Altss applies this (Altss methodology)

Altss methodology uses ordinal confidence labels per claim, not a numeric score. The labels (HIGH, MODERATE, LOW, SUSPENDED and UNASSESSED), which Altss calls Source Confidence, and the rules that assign them are Altss-defined and set out in the [Source Reliability & Confidence Methodology](https://altss.com/knowledge-center/frameworks/source-reliability-and-confidence-methodology). SUSPENDED applies while a claim's validation status is CONFLICTING: no strength is stated until a review resolves the conflict.

The label is kept separate from the three evidence dimensions it draws on. **Evidence origin** records what kind of source stated the claim; **derivation status** records whether the value was OBSERVED, DERIVED or ESTIMATED; **validation status** (UNVERIFIED, CORROBORATED, RESEARCH_VALIDATED or CONFLICTING) records corroboration and review, with a date. A label summarises how strong that evidence base was on the assessment date. It is not a probability that the claim is true, it does not describe the whole record, it is not an assessment of the entity, and it does not replace any of the three fields. A review raises a label only through evidence: a new independent source, a corrected provenance record, or a recorded check of the claim against its evidence. A reviewer's opinion alone does not, and any departure from the rules is recorded as an override. No calibration figure is implied by this description.

## Not the same as

- [Source Reliability](https://altss.com/glossary/source-reliability): Reliability grades a source; a confidence score is about a specific claim, match or judgment, with reliability as one input.

- [Precision and Recall](https://altss.com/glossary/precision-and-recall): Precision and recall measure how accurate a matcher's decisions were against checked outcomes; a confidence score is the matcher's own statement of support, which may or may not be calibrated.

- [Triangulation (Source Triangulation)](https://altss.com/glossary/source-triangulation): Triangulation checks whether a claim is supported by independent sources; a confidence score summarises the strength of support, of which corroboration is one input.

## Common mistakes

- Reading an ordinal label or an uncalibrated model score as a probability that the claim is true.

- Combining likelihood and confidence in one statement, such as "high confidence it is likely".

- Comparing scores from providers or models that use different methods.

- Averaging claim-level scores into one score for a record or an entity.

- Treating a high score as evidence that the fact is current.

- Confusing a confidence score with a statistical confidence level or confidence interval, which describe sampling uncertainty around an estimate.

- Raising a score because a reviewer agrees with the claim, without adding evidence.

## Edge cases

- A score for an absence ("no Form D found") is only as strong as the search behind it.

- A matching model can give a high score to two records that a register shows to be distinct entities; the register evidence prevails whatever the score says.

- For an estimated value, confidence describes the method and inputs; a stated range carries more information than a label.

- A derived value cannot be better supported than its weakest input.

## Questions

### Is a 90% confidence score a 90% chance of being right?

Only if the score has been calibrated: checked against independently verified outcomes for a stated population and period. Otherwise it ranks claims but does not measure accuracy.

### What is the difference between confidence and likelihood?

Likelihood is how probable an event or claim is. Confidence is how strong the evidence and reasoning behind that assessment are. ICD 203 requires the two to be stated separately.

## External standards

| Standard | Relation | Note |
| --- | --- | --- |

| ICD 203 Analytic Standards (US) (Tradecraft standard D.6.e(2), (2)(a) likelihood terms, (2)(b) confidence vs likelihood) | related |  |

## Sources

- [Intelligence Community Directive 203: Analytic Standards](https://archive.dni.gov/files/documents/ICD/ICD-203.pdf). Office of the Director of National Intelligence, ODNI, Signed 2 January 2015; technical amendments incl. 21 January 2022. Status: In force (as amended) (checked 2026-10-01). Section D.6.e(2), including (2)(a) table of likelihood terms and (2)(b) — supports: Likelihood vs confidence; basis of confidence; standard likelihood terms and ranges; prohibition on combining the two in one sentence

- [Improving information evaluation for intelligence production](https://doi.org/10.1080/02684527.2019.1569343). Daniel Irwin; David R. Mandel, Intelligence and National Security, Vol. 34(4), pp. 503-525. Status: Published (paywalled) (checked 2026-10-01). Abstract; Vol. 34(4), pp. 503-525 — supports: Proposes numeric probability estimates of information accuracy and periodic re-evaluation in place of rigid information evaluation methods

- [A Theory for Record Linkage](https://doi.org/10.1080/01621459.1969.10501049). Ivan P. Fellegi; Alan B. Sunter, Journal of the American Statistical Association, Vol. 64(328), pp. 1183-1210. Status: Published (paywalled) (checked 2026-10-01). Vol. 64(328), pp. 1183-1210 — supports: Match weights from agreement patterns as an example of a model score with decision thresholds

## Related terms

6 terms

- [Source Reliability](https://altss.com/glossary/source-reliability)

- [Triangulation (Source Triangulation)](https://altss.com/glossary/source-triangulation)

- [Precision and Recall](https://altss.com/glossary/precision-and-recall)

- [Entity Resolution](https://altss.com/glossary/entity-resolution)

- [Data Provenance](https://altss.com/glossary/data-provenance)

- [Data Freshness](https://altss.com/glossary/data-freshness)

## Concept record

Concept ID

ALTSS-DATA-020

Classification

Evidence & data

Topics

Private markets data & OSINT

Version

2.0.0

Last reviewed

2026-10-01

Structured data

[JSON](https://altss.com/reference/concepts/confidence-score.json)

## Canonical URL

https://altss.com/glossary/confidence-score
