Context length, throughput and latency: measuring generation
Four numbers describe a generation; none of them means anything without its conditions.

The L1 sheet in this area explains what a context window is and where its practical limits lie. This sheet is about the numbers people quote when they say a model is fast, and about the conditions without which those numbers mean nothing. There are four of them, they measure different things, and two of them move in opposite directions on purpose.
The four numbers
Context length — the number of tokens one request may occupy, prompt and reserved output together. It
is a setting as much as a model property. A llama.cpp server takes -c, --ctx-size, whose default of
0 means “loaded from model”, and it serves several requests at once through -np, --parallel server
slots which share the key-value pool; --kv-unified-per-slot N sets the “context limit per parallel slot”.
“The model supports 128k” and “this deployment serves 128k” are therefore different claims, and the second
one is the one that decides whether your request is truncated.
Time to first token (TTFT) — “How quickly users start seeing the model’s output after entering their query”, “driven by the time required to process the prompt and then generate the first output token”. It is the prompt-processing cost plus whatever queue the request waited in. A serving engine’s own metric for it is measured from the moment the request arrived, “in order to account for input processing time” — a reminder that a client-side measure includes the network and the queue as well as the model.
Time per output token (TPOT), also called inter-token latency — “Time to generate an output token for
each user that is querying our system”, which is the number a reader perceives as speed: “a TPOT of
100 milliseconds/tok would be 10 tokens per second per user”. A serving engine exports it under that name
(time_per_output_token_seconds, described as “Inter-token latency (Time Per Output Token, TPOT)”).
Throughput — “The number of output tokens per second an inference server can generate across all users and requests”. This is a server-level number, and it is the only one of the four that improves when more requests are served at once.
A serving engine publishes all of them as histograms — time to first token, time per output token, end-to-end request latency and queue time — so that the distribution, not just a mean, is visible.
The trade-off that makes single numbers useless
“If we process 16 user queries concurrently, we’ll have higher throughput compared to running the queries sequentially, but we’ll take longer to generate output tokens for each user.” That is not a defect to be tuned away: it is the same bandwidth bottleneck as the previous sheet, seen from the operator’s side. One pass over the weights can serve several sequences at once, so batching converts latency into throughput.
For a user-facing estimate of total latency the same source gives a rule of thumb: “Output length dominates overall response latency: for average latency, you can usually just take your expected/max output token length and multiply it by an overall average time per output token for the model.”
Which means: a throughput figure quoted alone is not a performance claim, it is a configuration.
Why the context length changes all four
- TTFT grows with the prompt. The prompt has to be processed before the first token exists, so a long retrieved document is paid for up front — before any output.
- The cache grows with the context, which is why a context limit is also a memory decision, and why a deployment can be fast at one context length and refuse work at another.
- The usable part of a long context is not uniform. A window that accepts 128k tokens is not a promise that a fact placed anywhere inside it will be used; the L1 sheet covers why the practical ceiling is lower than the advertised one.
How to report a number so that it means something
The area’s own rule is that performance figures are reported with model, runtime, hardware, version and measurement method — and to those this sheet adds what the numbers above make necessary: prompt length, output length, context limit and concurrency. A tokens-per-second figure without them cannot be compared with anything, including a later figure from the same machine.
Two failure modes are worth naming, because both look like results:
- A throughput figure in a chat. A single interactive request measures TPOT, not throughput; quoting it as throughput inflates the number by the batch size that was never used.
- A time-to-first-token figure that includes the cold start. The first request after a model is loaded
pays for loading it, which in a llama.cpp server is why
--timeoutexists and why a model kept warm behaves differently from one loaded on demand.
Level and prerequisites. L2 — operational: the reader must be able to name the four quantities, say which one a claim is about, and report a measurement with its conditions. No benchmark is run here and no target is set; measuring a specific deployment belongs to Progetti. Prerequisites: the L1 sheet on the context window, and the memory and placement sheets of this batch.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Model behaviour — the node this sheet sits in.
References
- Databricks — LLM Inference Performance Engineering — the definitions of TTFT, TPOT and throughput, the batching trade-off and the output-length rule of thumb.
- vLLM — metrics documentation — the metric names and the definitions the engine exports, including the arrival-time basis of the first-token metric.
- ggml-org — llama.cpp server —
--ctx-size,--parallel,--kv-unified-per-slotand the server timeout.