Temperature, top-p and output limits
How one token is chosen out of a distribution, and what a token limit actually bounds.

A language model does not decide on a sentence and then type it out. At each step it produces a probability distribution over its vocabulary for the next token, and a separate step — decoding — picks one token from that distribution. Temperature, top-p (nucleus sampling) and top-k act on that step: they reshape and truncate the distribution before a token is taken. They add nothing to what the model knows and they check nothing. An output-token limit is a different control — a ceiling on how much may be generated, not a length the model aims for.
The mental model: a distribution, then a choice
The model’s job and the decoder’s job are two different jobs, and the order matters:
context tokens ──► model ──► probability for every token in the vocabulary
│
decoding reshapes and truncates the distribution
temperature · top-p (nucleus) · top-k
│
keep a candidate set, discard the tail
│
greedy: take the maximum | sampling: draw
│
one chosen token
│
append it to the context, repeat (autoregressive)
▼
stop at end-of-sequence, or at the output limit
The model turns the context into a distribution; the parameters reshape and truncate it; one token is drawn from what is left and appended to the context, and the loop runs again. Everything to the right of the distribution is arithmetic over probabilities — no lookup, no check against the world.
Terminology the reader needs
- Decoding — the algorithm that turns the model’s per-step distribution into tokens; the Hugging Face guide names three families: greedy search, beam search and sampling.
- Greedy search — “the simplest decoding method”; it “selects the word with the highest probability as its next word” at each step. Beam search keeps several candidates at once but is “not guaranteed to find the most likely output”.
- Sampling — “randomly picking the next word … according to its conditional probability distribution”; generation this way “is not deterministic anymore”.
- Temperature — OpenAI: “sampling temperature … between 0 and 2”, default 1.0 in the examples; 0.8 makes output “more random”, 0.2 makes it “more focused and deterministic”.
- Top-p (nucleus sampling) — OpenAI: “an alternative to sampling with temperature”; only “the tokens with top_p probability mass” are considered, so 0.1 means the top 10% probability mass.
- Top-k — from the Hugging Face guide: “the K most likely next words are filtered and the probability
mass is redistributed among only those K next words”;
top_k=0deactivates it. - Output limit — OpenAI’s
max_completion_tokens: “an upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens”.
The mechanism, step by step
- The model reads the context and returns, for the next position, a probability over every token in its vocabulary; it is autoregressive, and what a token is belongs to a sibling sheet.
- The candidate set can be narrowed. Top-p keeps the smallest set of words whose cumulative probability exceeds p and redistributes the mass among them; top-k keeps the K most likely and does the same. Neither source read here states an order between this and temperature, and the two are separate operations on the same distribution.
- Temperature changes the shape of the distribution. Lowering it sharpens it — increasing high-probability likelihoods and decreasing low-probability ones; raising it flattens it.
- A token is chosen. Greedy takes the most probable; sampling draws one at random from the reshaped, truncated distribution.
- The token joins the context and the loop repeats, ending at an end-of-sequence token or at the output limit.
- On combining top-p with temperature, OpenAI’s note is: “We generally recommend altering this or temperature but not both.”
Greedy, temperature and top-p compared
| Approach | What it does | Parameter (OpenAI spec unless noted) | What the reader should expect |
|---|---|---|---|
| Greedy decoding | takes the most probable next token at each step | none — the guide describes it as the simplest method | as the guide puts it, it can miss high-probability words hidden behind a low-probability one |
| Temperature sampling | draws from the distribution after temperature has reshaped it | temperature, between 0 and 2, default 1.0 in the examples |
lower values narrow the distribution, higher values spread it out |
| Top-p (nucleus) | samples only from the smallest set of tokens whose cumulative probability exceeds p | top_p, 0 to 1, default 1.0 in the examples |
the size of the candidate set adapts to how predictable the next token is |
| Top-k | samples only from the K most likely tokens | top_k in the Hugging Face guide; top_k=0 deactivates it |
a fixed-size candidate set, whatever the shape of the distribution |
The output limit: a ceiling, not a target
max_completion_tokens is documented as “an upper bound for the number of tokens that can be generated for a
completion, including visible output tokens and reasoning tokens”. Two consequences follow. First, the bound
counts everything the model generates, including reasoning tokens the caller may never see, so the visible
answer can be shorter than the limit suggests. Second, an upper bound stops generation at the cap; it does not
make the model conclude. Read that as my framing: the specification defines a bound, and the difference
between “stopped because it finished” and “stopped because it hit the ceiling” is what a caller has to handle.
The context window is a different limit, owned by another sheet.
Limits, and the common conceptual errors
- Temperature is not a correctness dial. The specification’s wording is about random versus focused and deterministic — not about truth; neither source read here says a low temperature makes output more accurate, and this sheet claims no such thing. Why fluent output can still be false belongs to the hallucination and verification sheet.
- Temperature 0 is not a promise of identical output. The guide says that in its limit, as temperature
approaches 0, temperature-scaled sampling “becomes equal to greedy decoding”; the specification’s
seedparameter says “our system will make a best effort to sample deterministically”, and “Determinism is not guaranteed”. - Sampling introduces variation on purpose. With sampling, the guide says, generation “is not deterministic anymore” — the mechanism behind two identical prompts producing two different answers.
- A limit is not a length instruction. Asking for a long answer under a small cap yields text cut at the cap: it truncates, it does not summarise.
Level and prerequisites
L1 — the vocabulary of decoding and the difference between shaping a choice and bounding a length; no procedure and no tuning advice. Prerequisites: the sheets on tokens and autoregressive generation, and on prompts and messages; neither is published yet.
Where to go next
- Automation & AI — the area this sheet belongs to.
References
- OpenAI, OpenAI API specification (
openai-openapi.yaml) — the documented descriptions oftemperature,top_p,max_completion_tokensandseed, and the note about altering temperature or top_p but not both. - Hugging Face, How to generate text: using different decoding methods for language generation with
Transformers (
hf-how-to-generate.html, edited July 2023) — the three decoding families, greedy search, beam search, sampling, temperature, top-k and top-p (nucleus) sampling.