What a model knows: embedded knowledge and its limits
What sits in a model's parameters, what its training cutoff freezes, and what you supply.

A language model’s knowledge sits in its parameters — the trained weights fixed when training ends. They encode statistical patterns of the training data rather than an addressable record of it; verbatim memorisation is the documented exception, and it belongs to the sheet on privacy and data disclosure risk. Once training ends the parameters do not change: the model’s picture of the world stops at its training-data cutoff. It has no connection to your systems, your documents, your monitoring or today’s events. So whatever the model must reason about from your world has to enter the context of the call.
The mental model: what the model has, and what you must supply
what the model already has what you must supply
-------------------------- ---------------------
parameters: patterns learned from context: the input of this call
a text corpus, frozen at a cutoff |
| +-- your documents and data
| +-- retrieval results (L3)
| +-- tool results (L3)
v v
+-------------------------------------+
| model: predicts the next token |
+-------------------------------------+
no access to your systems, your files,
your monitoring, or today's events
Two supplies, from two different origins. The parameters are what the model arrived with, and they are static. The context is what you hand it in this call, and it is the only place current, private or local facts can enter.
Terminology the reader needs
- Parameters — the trained weights of the model. NIST describes generative models as producing outputs
that “approximate the statistical distribution of their training data” (
NIST.AI.600-1lines 306–308); the parameters are where that distribution is encoded. - Training-data cutoff — the point past which the training material does not run. Some model documentation
publishes it per model: the model overview page carries a “Training data cutoff” and a “Reliable knowledge
cutoff” row (
anthropic-models-overview.html). The dates are deliberately not repeated here; read them from the page for your model. - Parametric knowledge — what the model reproduces from its parameters alone, with no outside input.
- Context — everything passed in the request. The context-window sheet owns the window itself; here it matters only as the delivery channel for facts the model does not hold.
Why parameters are patterns, not records
NIST states the mechanism directly: confabulations “are a natural result of the way generative models are
designed: they generate outputs that approximate the statistical distribution of their training data”
(NIST.AI.600-1 lines 306–308). A database returns the row it stored, or returns nothing; a model returns
the most plausible continuation of the patterns it learned. Picking a word by its likelihood is not the same
operation as looking up an entry.
My own framing, not a sourced claim: parametric knowledge is not a database. A database has a schema, exact lookup and a way to return “no row”; parameters have none of these — they compress many observations into weights that favour typical continuations.
What the cutoff freezes, and what the model cannot know
Training ends on a date. Everything after it — a new release, an incident this morning, a change in your own
infrastructure — is outside the parameters and stays outside until a later model is trained. The AI RMF
defines an AI system as a system that, “for a given set of objectives, generate outputs such as
predictions, recommendations, or decisions” (NIST.AI.100-1 lines 174–177). My framing, labelled as such:
nothing in that definition gives the system a channel to read outside its inputs, and none exists unless the
application builds one.
Which questions the parameters can answer, and which need context
| Question | From the parameters alone | Needs to be supplied in context |
|---|---|---|
| A widely written-about general concept | usually plausible: well represented in training text | — |
| The contents of your internal runbook | not available | the document itself |
| An event from this morning | not available, and not until a later model is trained | a retrieval result or a tool call |
| A rare, contested or recently changed fact | the most frequent human answer, which may be the common misconception | a primary source, supplied |
Limits, and the common conceptual error
The error is to read fluency as coverage: a model answering outside its knowledge does not go silent, so a
gap in its knowledge and a wrong belief are indistinguishable in the output. NIST notes that users believe
false content “often due to the confident nature of the response” (NIST.AI.600-1 line 314).
TruthfulQA reports that models “often mimic popular misconceptions” (line 98) and that “the largest models
were generally the least truthful” (lines 25–26), naming answers that are “imitative falsehoods” because they
have high likelihood on the training distribution (lines 44–46). Size is no evidence that the parameters
hold what you need.
Model cards point the same way: a model should ship with its intended use cases and caveats, so users
can see the scope it was built for and avoid “contexts for which they are not well suited”
(model-cards-1810.03993.txt lines 13–15; sections “Intended Use” and “Caveats and Recommendations”, lines
235 and 389). What validity and reliability mean for an AI system — and what a check has to compare
against — belongs to the sibling sheet on hallucination, uncertainty and verification.
The consequence is the point of this sheet: put the data in the context. Retrieval — finding relevant documents and inserting them into the request — is one way; tool use — letting the model call something that returns live data — is another. Both are L3 subjects, named here only.
What to remember
- Parameters hold statistical patterns rather than an addressable record of the training data.
- The cutoff freezes the model’s world until a later model is trained.
- Absence and wrong knowledge look the same in the output, because both arrive as fluent text.
- Supply the data in the context; retrieval and tool use are the standard ways, developed at L3.
Level and prerequisites
L1 — the model’s embedded knowledge and its limits, with no procedure, no retrieval configuration and no tuning recipe. Prerequisites: none beyond reading the Automation & AI area.
Where to go next
- Automation & AI — the area this sheet belongs to.
References
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023) — the definition of an AI system as a system generating outputs for a given set of objectives, and the trustworthy-AI characteristic “valid and reliable” to be established through measurement.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024) — the statement that generative models approximate the statistical distribution of their training data, and that confident responses lead users to believe false content.
- S. Lin, J. Hilton, O. Evans, TruthfulQA: Measuring How Models Mimic Human Falsehoods (2021) — imitative falsehoods, popular misconceptions, and the inverse-scaling finding that larger models are not automatically more truthful.
- M. Mitchell et al., Model Cards for Model Reporting (2019) — the practice of documenting a model’s intended use cases and its caveats.
- Anthropic, Models overview (vendor documentation) — the published per-model training-data cutoff field (and the “Reliable knowledge cutoff” field), cited for the existence of the field, not for any date.