Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Error codes

Every error LUMEN returns to a client carries a stable LM-XXXX code, an HTTP status, and a coarse type. The response body is always:

{ "error": { "code": "LM-1001", "message": "…", "type": "invalid_request" } }

The type is one of invalid_request, upstream_error, internal, or client_cancelled. The gateway always distinguishes these situations and never disguises one as another - in particular, an internal malfunction is never reported as a misleading 401 (a lesson from OpenRouter outages), a malformed upstream response is a 502, never a gateway 500, and a client-initiated cancel (client_cancelled) is never counted as an internal malfunction either (see LM-6xxx below).

Codes are stable: once assigned, a code keeps its meaning across releases. The code prefix groups by cause: 1xxx request, 2xxx routing, 3xxx upstream, 4xxx auth/budget, 5xxx internal, 6xxx client-cancellation.

Request errors - LM-1xxx · type: invalid_request

CodeHTTPMeaning
LM-1001400Malformed or invalid request body / parameters. Also returned when a provider in strict mode is sent an unsupported-but-meaningful field it cannot honor (e.g. dimensions to Ollama, or response_format/seed/logprobs/parallel_tool_calls to a translated chat provider that has no native equivalent - see the chat-extras matrix), and when an input shape a provider cannot consume at all is sent (pre-tokenized token-id arrays to a text-only embed API: Cohere, TEI, Ollama, Jina, Voyage, Mistral). Both are rejected before any upstream call; the message names the field/shape and provider.
LM-1002413Request body exceeded the configured size limit.
LM-1003404No route matches the request method and path. Returned by the router fallback for trailing-slash, extra-segment and other near-miss paths, so an unmatched request carries the same envelope as every other rejection instead of a bare, empty-body 404. The message names no path and discloses no route. The fallback sits outside the virtual-key auth layer, so an unmatched path answers 404 even when auth is enabled (it leaks no more than the bare 404 did); a matched /v1 route without a key is still LM-4004. Distinct from LM-2001, which means the HTTP route matched but the requested model id does not exist.
LM-1004412PUT /admin/config: the submitted If-Match no longer matches the config file’s current hash (ADR 010) - the file changed since it was read, most often another operator’s apply landing first. The request is refused before anything is staged or written; re-GET /admin/config for the current hash and content, then re-apply. Kept distinct from a missing/malformed If-Match header, which is a plain LM-1001 400: the console needs to tell “you sent no token” from “your token is stale” apart, since only the second case calls for a re-read-and-retry rather than a client bug fix.

Routing & capability-request errors - LM-2xxx · type: invalid_request

CodeHTTPMeaning
LM-2001404The requested model id was not found.
LM-2002400The model exists but does not serve the requested capability.
LM-2003400An image content part was sent to a model without the image modality (chat vision M8 and embeddings M9).
LM-2004400A remote image URL was sent to a provider that only accepts inline base64 image data (chat vision M8). Checked before any upstream call for the primary route; if a fail-over reaches an image-incapable fallback further down the chain, the same code and status surface there too, naming the fallback provider - never the generic LM-3002 a translation failure would otherwise produce (GH #13).
LM-2005400A remote image URL was supplied to /v1/embeddings but server-side image fetching is disabled ([image_fetch] enabled = false). Inline the image as a data: URI or enable fetching (M9).
LM-2006400A remote image URL was rejected by a fetch guard (scheme, host/prefix allowlist, private-IP block, size cap, per-request count cap, or non-image content type). The specific reason is logged server-side, never returned (M9).
LM-2007502A permitted image fetch failed at the remote host (network error, timeout, or error status). type: upstream_error (M9).
LM-2008400A provider-native image source (Anthropic file_id, spelled anthropic-file:<id>; Gemini fileUri, a gs:// GCS URI or a Gemini Files API URI) was sent to a provider that cannot resolve it - the resolved primary provider must match the reference’s own provider.
LM-2010400A rerank request supplied no documents to score.

Upstream errors - LM-3xxx · type: upstream_error

These always name the provider that failed. Retriable ones may be transparently retried on a fallback before surfacing.

CodeHTTPMeaning
LM-3001429An upstream provider rate limited the request.
LM-3002502An upstream provider returned an unparseable/malformed response.
LM-3003502An upstream provider returned an error status.
LM-3004503No healthy upstream available (circuit open / fallbacks spent).
LM-3005504An upstream provider timed out.
LM-3010502An upstream stream ended prematurely (no terminator).
LM-3011504An upstream produced no first token within the first-token deadline.
LM-3012504The connection to an upstream could not be established within the connect timeout.
LM-3013504The whole request (all retries + fallbacks) exceeded the total timeout.
LM-3020503The provider’s circuit breaker is open and no fallback remained.

For LM-3001, LM-3020 (and LM-4002/LM-4003), a Retry-After value may be advertised. The three timeouts (LM-3011 first-token, LM-3012 connect, LM-3013 total) are distinct codes purely for debugging - see §6.4 and docs/adr/005-resilience-execution.md.

How resilience shapes these codes

The 3xxx codes are what a client sees only after the resilience machinery has given up. Before surfacing, a retryable failure (LM-3001 429, LM-3003 5xx, LM-3005/LM-3012 timeouts) is retried with exponential backoff, then the request fails over to the model’s configured fallbacks. The mapping between a failure and the code that eventually surfaces:

  • LM-3020 (503) - the primary’s circuit is open and no fallback remained. Skipping an open circuit is instant (no upstream call), and the response carries a Retry-After equal to the cooldown remainder.
  • LM-3004 (503) - every link in the fallback chain was tried and failed (retries exhausted or circuits open all the way down).
  • LM-3013 (504) - the total per-request deadline elapsed while retrying or failing over; it bounds all attempts together, so a slow chain fails here rather than hanging.
  • LM-3011 / LM-3012 (504) - first-token and connect timeouts; each is a retryable failure on its own before it surfaces.

A hard upstream client error (a 4xx bad request) is never retried or failed over - a different provider would reject it too - and surfaces immediately. Whichever model ultimately served a successful request is reported in the x-lumen-model-used response header.

Auth / budget errors - LM-4xxx · type: invalid_request

Codes pinned by the spec. Enforcement happens in memory, before any upstream call - a rejected request never leaks spend to a provider.

CodeHTTPMeaning
LM-4001402A hard budget the key is subject to is exhausted: either the key’s own budget_max or its budget group’s shared pool (ADR 009). Same code and status for both; the message text discloses the scope (“budget exceeded for this key” vs “budget exceeded for this key’s group”).
LM-4002429The key’s requests-per-minute quota was exceeded.
LM-4003429The key’s tokens-per-minute quota was exceeded.
LM-4004401Missing or invalid virtual key. Deliberately does not say why (unknown, disabled and expired are indistinguishable) so callers cannot probe key state.

Internal errors - LM-5xxx · type: internal

CodeHTTPMeaning
LM-5001500Internal gateway malfunction.

Internal errors return an opaque "internal error" message to the client; the underlying detail is written only to the server logs, never the response.

Client-cancellation - LM-6xxx · type: client_cancelled

CodeHTTPMeaning
LM-6001499The client disconnected before the request completed; the upstream call was aborted.

499 is the conventional “client closed request” status (nginx). The client is normally already gone by the time this would be returned, so the status exists for logs and metrics, not for anything a client reads. It is deliberately kept out of both type: internal and the 5xx status class: a client hanging up is not a gateway malfunction, and must never inflate the internal-error metrics or alerts a real one would (issue #11, see docs/adr/006-client-cancellation-error-code.md).

Two paths produce it. A cancellation surfacing mid-stream is emitted as a terminal SSE error frame carrying this envelope. A client that simply disconnects mid-stream never sees a frame at all, but the request’s accounting record (usage_log.status and the lumen_request_duration_seconds{status="499"} sample) is settled at 499 instead of being miscounted as a 200 success. A non-streaming disconnect drops the request before any outcome is recorded and produces no sample.