Model serving and compatible APIs: one endpoint, several contracts
A serving engine adds a queue, a lifecycle and an API surface to a weights file.

A weights file cannot answer a request. Something has to hold the model in memory, decide how many requests to run at once, queue the rest, and speak HTTP — that something is the serving engine, and it is where the phrase “OpenAI-compatible” does its work. The habit worth building is this: treat compatibility as a claim about shape, verify it against the engine’s own documentation, and remember that the same engine usually offers more than one contract at the same time.
What serving adds to a model
A model file plus a server is a different kind of object from a model file:
- residency — the model stays loaded between requests; in Ollama the
keep_aliveparameter controls how long it stays in memory after the last request. - a queue and a batching strategy — server slots (
-np, --parallel) and continuous batching (-cb, --cont-batching, enabled by default), which is what turns several requests into one batch. - a lifecycle — which model is loaded, when it is loaded, and what happens to the first request that arrives before it is ready.
- an HTTP surface — endpoints, authentication, streaming, error shapes and, optionally, metrics.
The surface a local engine actually exposes
The llama.cpp server is a useful concrete case because its documentation lists every route. Its defaults
are deliberately conservative: --host defaults to 127.0.0.1 and --port to 8080.
A health endpoint that needs no key. GET /health is documented as “public (no API key check)” with
/v1/health as an alias, and it reports 503 while the model is not ready — which is precisely what a
load balancer or a start-up script wants to poll.
OpenAI-compatible routes. GET /v1/models (“OpenAI-compatible Model Info API”), POST /v1/completions, POST /v1/chat/completions, POST /v1/embeddings and POST /v1/responses are each
documented as OpenAI-compatible APIs.
A second contract. POST /v1/messages is documented as an “Anthropic-compatible Messages API” — the
same server, the same loaded model, a different request and response shape.
Native routes. Alongside them the server exposes its own endpoints: /completion, /tokenize,
/detokenize, /apply-template, /infill, /props, /slots, /lora-adapters. They are richer than the
compatible ones — and, being engine-specific, they are not portable.
Observability and security. GET /metrics is a “Prometheus compatible metrics exporter”, disabled
until --metrics is passed; authentication is off until you pass --api-key (a comma-separated list) or
--api-key-file; TLS is optional (--ssl-key-file, --ssl-cert-file).
What “compatible” does and does not mean
It means the route and the JSON shape match closely enough that a client written for the OpenAI API works
unchanged. The engine’s own documentation marks the seam explicitly: some endpoints are “not
OAI-compatible” — /completion and /embeddings are described that way, next to their OpenAI-compatible
counterparts /v1/completions and /v1/embeddings, which exist precisely because the response formats
differ.
What the shape does not cover: default values for parameters the caller omits, fields the local engine adds to a response, the exact body and status of an error, rate-limit headers (a local engine has none in the provider’s sense — see the API sheet in this batch), model naming and aliases, and whether a parameter is implemented at all. Two engines can both be “compatible” and disagree about every one of those.
The consequence of running one
A local server is an HTTP service that, by default, listens only on loopback and accepts any caller. Making
it reachable from the network is a second decision, and it is separate from the first: --host 0.0.0.0
publishes the port, while --api-key is what decides who may use it. Publishing without a key exposes the
machine’s accelerator, its loaded model and its logs to anyone who can reach the port. The /health
endpoint stays public by design, which is reasonable because it discloses only readiness — and it is worth
knowing when reading an engine’s threat surface.
Level and prerequisites. L2 — operational: the reader must be able to bring a model up behind an HTTP interface, name what the serving layer adds, and check a compatibility claim instead of trusting it. No server is installed or configured in this sheet, and no deployment is described. Prerequisites: the memory and placement sheets of this batch.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Applications — the node this sheet sits in.
References
- ggml-org — llama.cpp server — the default host and port, the full endpoint list with each route’s own description, the API-key options, TLS options and the metrics endpoint.
- Ollama — API reference — a second local engine’s surface:
POST /api/generateandPOST /api/chatonlocalhost:11434, with streaming by default. - vLLM — documentation — a serving engine described as exposing an “OpenAI-compatible API server, plus Anthropic Messages API and gRPC support”.