Structured JSON output and schema validation
Asking for JSON and constraining the decoder to JSON are two different guarantees.

“Give me JSON” and “only this JSON” are different requests, and the difference is not politeness. The first is an instruction the model may follow; the second is a constraint on which tokens the decoder is allowed to produce at all. Engines document them as separate options — the Ollama API reference, for instance, has one section for structured outputs and another for JSON mode — and treating them as variants of the same feature is how a pipeline ends up parsing text that is almost JSON in production.
Two mechanisms
Prompted JSON. You ask for JSON in the prompt. The model usually complies, and nothing in the sampling process prevents a missing brace, a sentence before the object, or an invented value for a field you described as one of three choices. This is a request, and the failure mode is malformed or schema-violating text.
Constrained decoding. The sampler is restricted to the language a grammar defines, so the text cannot
leave it. In a llama.cpp server this is --grammar GRAMMAR, documented as a “BNF-like grammar to constrain
generations”, with a per-request form: the json_schema field, documented as “Set a JSON schema for
grammar-based sampling (e.g. {"items": {"type": "string"}, "minItems": 10, "maxItems": 100} of a list of
strings, or {} for any JSON)”. The hosted-service version of the same idea is Structured Outputs, whose
purpose is stated as “Ensure text responses from the model adhere to a JSON schema you define”, with the
promise made explicit: the feature “ensures the model will always generate responses that adhere to your
supplied JSON Schema, so you don’t need to worry about the model omitting a required key, or hallucinating
an invalid enum value.”
Note what that sentence promises: shape, reliably. Nothing in it says the values are correct.
What a schema actually constrains
A JSON Schema is a document that makes assertions about what a valid instance looks like — the validation specification describes it as defining “assertions about what a valid document must look like”. Its keywords are conservative by default, which is the trap:
- The
propertieskeyword defines the properties an object has. Any property that does not match one of the names inpropertiesis ignored by this keyword — so an extra key is not an error unless something forbids it, and forbidding it is a separate decision (the specification points toadditionalPropertiesfor exactly that). - Requiring fields is also explicit: “By default, leaving out properties is valid.” A field is only
mandatory if
requiredsays so. - Constraining values is a third, separate decision:
type,enum, and the numeric and textual limits.
A schema can therefore look strict — it lists the fields you care about — and be permissive about everything else: extra keys, missing keys and the domain of each value are three independent settings.
One more thing is worth knowing before citing a schema as authority: the document read for this sheet is an Internet-Draft of the JSON Schema specification, and Internet-Drafts state about themselves that they are “working documents” and that “It is inappropriate to use Internet-Drafts as reference material”. The draft is useful as the technical definition of the keywords; it is not a standard you can cite at a customer, so speak of “JSON Schema, 2020-12” as a family of published drafts rather than as an RFC.
Validation is still yours
Constrained decoding removes one failure class and leaves the others intact:
- Shape is guaranteed, meaning is not. A schema-valid document can name a service that does not exist, quote a policy that was not in the retrieved text, or place a number in the wrong field of a well-typed object. Validation confirms the document is parseable and well-formed; it says nothing about truth.
- A refusal is not an answer. A model can decline the request rather than fill the schema, and a provider that makes this “programmatically detectable” is telling you to handle it as a distinct outcome — not to parse it as an empty result.
- The schema is part of the contract, so version it. A changed schema silently changes what the caller accepts; a pinned schema plus a validation step is what makes the pipeline reproducible.
- Log the raw text at the boundary, at least while a pipeline is being established. When validation fails, the raw output is the only evidence of what happened.
The practical sequence
- Write the schema so that “strict” means what you think: explicit
required, explicitadditionalProperties, explicit value domains. - Constrain decoding where the engine supports it, rather than asking politely.
- Validate the received document against the same schema before use, and fail loudly.
- Keep the refusal path and the raw text.
Level and prerequisites. L2 — operational: the reader must be able to choose between asking for JSON and constraining it, to write a schema that means what it says, and to keep validating. Prerequisites: the prompt-envelope sheet of this batch (how a request is assembled) and the L1 material on how models generate text.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Applications — the node this sheet sits in.
References
- JSON Schema — core and validation specifications (2020-12 drafts) — the definition of an assertion,
the
propertieskeyword and its permissive default, and the draft status of the documents. - JSON Schema — Understanding JSON Schema, object reference — “By default, leaving out properties is
valid” and the pointer to
additionalProperties. - OpenAI — Structured Outputs — what the feature guarantees, and refusals as a detectable outcome.
- ggml-org — llama.cpp server —
--grammarand the per-requestjson_schemafield. - Ollama — API reference — structured outputs and JSON mode documented as separate options.