Open Model Gatewaydocs

Errors and limits

Error codes and HTTP statuses of the inference API, which ones to retry, and the size limits.

Errors use the OpenAI shape, or Anthropic's on /v1/messages:

{"error": {"message": "Budget for this API key would be exceeded in its current period", "type": "insufficient_quota", "code": "budget_exceeded", "param": null}}

Messages are fixed strings. They never contain provider error text, amounts, identities or how much budget is left.

Codes

HTTPcodeMeaningRetry
400invalid_request_errorA field or value outside the supported subset, or malformed input.No: fix the request
400upstream_rejectedThe provider refused the request (validation, moderation, unknown model).No
400unsupported_capabilityImages, audio, rerank, System One, realtime, files, batches and videos: no route can serve this request.No
401Missing, repeated or invalid key, or a disabled, revoked or expired one.No
404model_not_foundNo such model, or not available to this key.No
404not_foundA file, batch or video id that doesn't exist or belongs to another workspace.No
413payload_too_large, file_too_large, storage_quota_exceededThe body, file or workspace storage is over its limit.No
429rate_limit_errorA per-minute or at-once limit, or the provider's rate limit.Yes, later
429job_limit_exceededA "jobs at once" limit (batches, videos).Yes, when a job ends
429token_reservation_exceeds_limitThe model's input plus output ceiling is larger than a tokens-per-minute limit.No
429budget_exceededA budget would be exceeded in its current window.No
429unresolved_usageUnknown, unbounded costs block budgeted requests.No
501unsupported_capabilityChat Completions, Responses, Messages and embeddings: no route supports the requested features.No
502upstream_unavailable, invalid_upstream_responseThe provider failed or answered with something invalid.Maybe
503provider_configuration_errorConfiguration problem: a missing credential, an unpriced meter, or a request without an output maximum on a priced model.No
503accounting_unavailableThe gateway couldn't record the request.Maybe
504timeout_errorThe request deadline passed (120 seconds unless the operator changed it).Maybe

Budget denials use the type insufficient_quota; token_reservation_exceeds_limit and job_limit_exceeded use rate_limit_error. On /v1/messages every 429 has Anthropic's rate_limit_error type, and the message says which case applies.

Don't retry what can't succeed

budget_exceeded, unresolved_usage and token_reservation_exceeds_limit responses carry x-should-retry: false. The official OpenAI and Anthropic SDKs honour it and don't retry. Errors from the Files API, except 503s, carry it too.

Which scope refused

A limit denial names the scope kind: this API key, this workspace, or the installation. When several refuse at once, the narrowest is reported (key, then workspace, then installation), and budget denials come before rate limits. The installation's budget always reports budget_exceeded, even if another workspace's unknown usage is the cause.

Size limits

Unless the operator changed them:

WhatLimit
JSON request body2 MiB
Image generation, rerank, System One, speech request body2 MiB each
Transcription upload25 MiB file, 26 MiB request
Files API upload200 MiB
Batch input50,000 lines, 4 MiB each
Provider response4 MiB (images: 20 MiB)
Embeddings128 inputs, 1 MiB of text in total
Request deadline120 seconds

The operator's settings are in Configuration.

On this page