Hallucination, uncertainty and verification
Why a fluent model answer can be false, and what checking it actually requires.

A language model produces a fluent, confident answer by the same route whether that answer is true or false. Nothing in the mechanism compares the claim with the world: outputs “approximate the statistical distribution of their training data” and predict “the next token or word in a sentence or phrase” (NIST AI 600-1, lines 306–310). NIST names the failure confabulation: “The production of confidently stated but erroneous or false content … by which users may be misled or deceived” (lines 204–205). Three consequences follow: fluency is not evidence; confidence in the text is a property of the generation, not a measurement of the claim; and verification means comparing the claim with a source outside the model.
The mental model: a claim, where it came from, and how it is checked
claim in the answer
│
├── from the parameters? → no source to point at
│ (learned at training, frozen afterwards)
│
└── from the supplied context? → the text is in front of you
│
▼
does the claim change a decision, a number or an action?
│
┌────────┴─────────┐
▼ ▼
REQUIRED check NOT REQUIRED
compare with a note it as unverified;
source outside do not rely on it
the model
Every factual assertion in an answer is a claim, and the useful first question is not “is it true?” but “where did it come from, and does it change a decision?”. Two origins matter: the model’s parameters (learned at training, then frozen) and the supplied context (the text you put into the conversation). A claim from the parameters cannot be pointed at a source; a claim from the supplied context can.
Terminology the reader needs
- Confabulation — confidently stated but erroneous or false content; “hallucination” and “fabrication” are the colloquial equivalents (lines 204–205, 237).
- Intrinsic hallucination — output that “contradicts the source content” (Ji et al., line 210): the source is in front of you and the statement disagrees with it.
- Extrinsic hallucination — output that “cannot be verified from the source content”, i.e. that “can neither be supported nor contradicted by the source” (lines 215–216); it “is not always erroneous” (line 220) but is treated with caution because it is unverifiable (lines 223–224).
- Imitative falsehood — a false statement learned from imitating human text; TruthfulQA’s term for answers that “mimic popular misconceptions” (Lin et al., lines 18, 98).
The mechanism, step by step
- Text is generated token by token from a probability distribution over what plausibly comes next, which reflects the training data; the same machinery “can produce factually accurate and consistent outputs” and “outputs that are factually inaccurate or internally inconsistent” (lines 306–310).
- Fluency survives error: confabulations include outputs that “diverge from the prompts or other input” or “contradict previously generated statements in the same context” (lines 302–305).
- The justification is generated the same way: outputs may “include confabulated logic or citations that purport to justify or explain the system’s answer”, listing steps “even when the answer itself is incorrect” (lines 321–324). Reasoning-shaped text is not an audit of the claim.
- Uncertainty in the answer is generated text too, not a calibrated signal: Ji et al. call teaching a model to “honestly admit ‘I don’t know’” crucial and unsolved, and state that of the methods tried, “none of them can satisfactorily solve this issue” (lines 2242, 2249). A hedge is a phrase, not a probability.
Confidence is not evidence: what a check must compare against
Two beliefs fail at the same point. The first is that a wrong answer would sound unsure: users believe false content “often due to the confident nature of the response” (line 314). The second is that a larger model is more truthful. TruthfulQA found that the largest models were generally the least truthful — the effect its authors call inverse scaling — and that imitative falsehoods are “not solved merely by scaling up” (lines 25–26, 98–100).
Verification is not a property of the answer; it is an act performed outside the model. NIST AI 100-1 lists “valid and reliable” among the characteristics of a trustworthy AI system — a claim about the system, not about any single response (lines 256, 590). In the framework’s own terms, checking is a documented, system-level activity: MEASURE 2.1 states that “Test sets, metrics, and details about the tools used during TEVV are documented” (line 1365), and the Generative AI Profile asks to verify the “sources and citations in GAI system outputs” (MS-2.5-003) and to document fact-checking of generative-AI output (MP-2.3-003). Neither asks whether the answer sounded confident.
Because a check costs something, the decision of which claims require one is made claim by claim, in advance:
| The claim in the answer | Where it came from | Check with a source outside the model | Reason |
|---|---|---|---|
| A statement about the world that the prompt does not contain | parameters (training, frozen) | REQUIRED | the model cannot point at a source, and nothing in the session can confirm it |
| A statement about a document or log you supplied | supplied context | REQUIRED | the source text is in front of you, so a contradiction is detectable — the intrinsic case |
| A plausible extra detail the supplied source never mentions | parameters, added alongside the context | REQUIRED, unless the text is deliberately creative | the source can neither support nor contradict it — the extrinsic case |
| A restatement of your instruction or of earlier context | supplied context | NOT REQUIRED | it repeats text you already hold; it is not a new claim about the world |
| Format, length, tone or phrasing of the answer | not a factual claim | NOT REQUIRED | there is no truth value to check |
Limits, and the common conceptual error
The common error is to read confidence as calibration and fluency as correctness: both are properties of a next-token process and are the same whichever way the answer turns out. A belief the sources do not support is worth naming too: that a lower sampling temperature, or a larger or newer model, removes confabulation. The Generative AI Profile locates the cause in how generative outputs are produced at all — statistical approximation of the training distribution (lines 306–310) — and neither source read here states that lower temperature or a larger model makes an answer truthful. The size half deserves separating: what TruthfulQA measured above is truthfulness rather than confabulation, and it points the other way for size. Treat that belief as unverified. Grounding an answer in retrieved sources is L3 material, named here only.
Level and prerequisites
L1 — vocabulary and mechanism: what confabulation is, why fluency and confidence are not evidence, and what a check compares against. No tool, no procedure, no configuration, no benchmark. Prerequisites: none.
Where to go next
- Automation & AI — the area this sheet belongs to.
References
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023), https://doi.org/10.6028/NIST.AI.100-1 — “valid and reliable” as a characteristic of trustworthy AI systems (lines 256, 590); MEASURE 2.1 on documented test sets, metrics and tools (line 1365).
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024), https://doi.org/10.6028/NIST.AI.600-1 — confabulation as a named risk (lines 204–205); its mechanism (lines 302–310); the confident nature of the response (line 314); confabulated logic and citations (lines 321–324); the output-verification measures MS-2.5-003 and MP-2.3-003.
- Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Chen, D., Dai, W., Chan, H. S., Madotto, A., Fung, P., Survey of Hallucination in Natural Language Generation, ACM Computing Surveys; arXiv:2202.03629v7 (14 July 2024) — intrinsic and extrinsic hallucination (lines 210, 215–216, 220, 223–224); expressing uncertainty as an unsolved problem (lines 2241–2250).
- Lin, S., Hilton, J., Evans, O., TruthfulQA: Measuring How Models Mimic Human Falsehoods, arXiv:2109.07958v2 (8 May 2022) — truthfulness as the benchmark’s object (line 11); false answers “learned from imitating human texts” (line 18); largest models generally least truthful and “inverse scaling” (lines 25–26, 98–100).