Articles

API authentication, rate limits and error handling

An endpoint refuses a request for three different reasons; the response says which.

Reading: 6 minAutomation & AI

Article cover: API authentication, rate limits and error handling

An LLM endpoint can refuse a request for three very different reasons — who you are, how much you are asking for, and whether the service is working — and the response tells you which one it is. Treating them alike, by retrying everything, is how a client turns one throttled request into a self-inflicted outage. The classification is the skill; the retry mechanics belong to the next sheet.

Authentication: who is asking

A provider API authenticates with a bearer credential. The OpenAI API description declares its security scheme as ApiKeyAuth of type: http and scheme: bearer, which is the standard way of saying that the caller must send Authorization: Bearer <credential>; a second scheme covers administrative operations at organisation level. Two failures are distinct and must be handled differently:

  • 401 — not authenticated. The credential is missing, malformed, revoked or wrong; Anthropic names this authentication_error in its own error table.
  • 403 — authenticated but not permitted. The key is valid and lacks the right; Anthropic names this permission_error.

A local engine inverts the default: authentication is something you turn on. In a llama.cpp server, --api-key KEY is documented with “default: none” — and the health endpoint is explicitly “public (no API key check)”. So an unauthenticated local server is not a misconfiguration that slipped through; it is what happens when nothing is configured, and the same documentation offers --ssl-key-file and --ssl-cert-file for TLS when the server leaves loopback. Where the credential lives is the subject of the secrets sheet in this batch.

Rate limits and quotas: how much, and how fast

The status code for “too much, too quickly” is 429, defined in its own specification section: it “indicates that the user has sent too many requests in a given amount of time (‘rate limiting’)”, and that the response “representations SHOULD include details explaining the condition, and MAY include a Retry-After header indicating how long to wait before making a new request”.

What a well-behaved client does with it is described in the provider’s own API description. Its rate-limited response documents a Retry-After header defined as “The minimum number of seconds to wait before retrying. This header is returned when the server has computed a retry delay and may be omitted” — with a minimum value of one second — and explains the condition in the body of the error: a slow_down error means traffic increased too quickly, so the client should reduce its request rate and then increase it gradually. “Gradually” is the operative word: a client that waits the stated delay and then resumes at exactly the rate that triggered the limit is asking for the same answer again.

Where the API documents them, per-endpoint rate-limit headers (X-RateLimit-Limit-*, X-RateLimit-Remaining-*, X-RateLimit-Reset-*) let a client observe its budget instead of discovering it by being refused.

429 is a message about you; a 5xx is a message about the service. Retrying a 429 immediately increases the load that caused it. A quota is not an outage.

A local engine has no quota in the provider’s sense, which does not mean it has no limits: it has a fixed number of server slots (-np, --parallel), a read/write timeout (-to, --timeout, default 3600 seconds) and a finite memory pool. Exceeding its capacity shows up as queueing or as a request that fails after waiting — not as a 429 with a retry hint.

The service’s own failures

The 5xx class is defined by the HTTP specification as failures the client can usually do nothing about: 500 Internal Server Error, 502 Bad Gateway, 503 Service Unavailable and 504 Gateway Timeout. On top of those, providers add codes of their own: Anthropic’s table lists 500 api_error and, distinctly, 529 overloaded_error — “the service is temporarily overloaded”, which in the core HTTP set has no exact equivalent. When a provider invents a status code, the only reliable instruction is that provider’s own documentation; a client that switches on the numeric range alone will misclassify it.

This is the class where retrying is reasonable — with backoff, a cap and jitter, which is the next sheet.

A classification you can implement

Response Meaning What to do
400, 422 the request is malformed fix the request; never retry unchanged
401 the credential is missing, invalid or revoked refresh or replace it; stop
403 the credential lacks the right escalate; retrying changes nothing
404 the route or the model identifier is wrong fix the reference
413 the payload is too large shorten the input
429 rate limit or quota honour Retry-After, then slow the rate
500, 502, 503, 504 the service failed retry with backoff, capped
529 (provider-specific) the service is overloaded as 5xx, per that provider’s documentation
timeout, no response unknown state the next sheet: idempotency decides whether a retry is safe

The last row is the one that matters most and the one most often ignored: a request that may have been executed is not the same as a request that failed.

The operational cost of getting this wrong

Retry logic is where a client’s behaviour becomes part of the service’s problem. OWASP names the class LLM10 Unbounded Consumption: excessive and uncontrolled inference, whether by design or by a client that multiplies its own traffic, producing denial of service, economic loss and service degradation. A retry policy is an amplifier; its controls — classification, backoff, cap, jitter, and a circuit breaker when a dependency is clearly down — are the subject of the next sheet.

Level and prerequisites. L2 — operational: the reader must be able to classify a refusal and choose a response policy for each class. Prerequisites: the serving sheet of this batch (what an endpoint is) and the secrets sheet for where the credential lives.

Where to go next

References

  • IETF — RFC 6585, §4 — the definition of 429 and the Retry-After guidance.
  • IETF — RFC 9110 — the definitions of 500, 502, 503 and 504, and the retry rule for non-idempotent methods.
  • OpenAI — API description — the bearer security scheme, the rate-limited response with its Retry-After header and its slow_down condition, and the per-endpoint rate-limit headers.
  • Anthropic — Errors — the error taxonomy including 401, 403, 413, 429, 500 and the provider-specific 529.
  • OWASP — Top 10 for LLM Applications (2025), LLM10 — unbounded consumption as a risk class.