Open Model Gatewaydocs

Chat Completions

POST /v1/chat/completions, the supported subset of OpenAI's Chat Completions API, streaming and per-provider notes.

POST /v1/chat/completions
{
  "model": "example/chat",
  "messages": [
    {"role": "system", "content": "You answer in one sentence."},
    {"role": "user", "content": "What is a reranker?"}
  ],
  "max_completion_tokens": 100,
  "stream": false
}

Request fields

FieldSupported
modelRequired: a gateway model name.
messagesRequired. Roles system, developer, user, assistant and tool. content is a string: content-part arrays (images, audio, files) are refused. Assistant tool_calls and tool tool_call_id are supported.
max_completion_tokensThe output maximum. Required in practice on a priced model, and no higher than its output ceiling. The legacy max_tokens is refused.
temperatureSupported (Anthropic: 0 to 1).
stream, stream_options.include_usageSupported.
tools, tool_choiceFunction tools only.
nAbsent or 1.
user, metadataAccepted as Logs labels only; never sent to the provider. metadata keeps OpenAI's bounds: up to 16 pairs, keys up to 64 and values up to 512 characters.

Any other field (for example response_format, reasoning_effort, logprobs, seed, top_p, stop, parallel_tool_calls) is refused with 400 invalid_request_error.

Response

The OpenAI response shape, with one choice, the gateway model name and usage with the token counts the provider reported. A count the provider didn't report is left out, never sent as 0.

Streaming

With "stream": true the response is server-sent events, delivered token by token. With stream_options.include_usage, the final chunk carries usage. A successful stream ends with data: [DONE]; an error after the stream started closes it without [DONE], never as a fake success.

Providers

ConnectionNotes
OpenAINative Chat Completions.
AnthropicTranslated to Messages: leading system/developer text, temperature up to 1; strict tool schemas and non-object tool arguments are refused.
Amazon BedrockTranslated to Converse: leading system messages only, no developer messages or assistant prefill; tool choice none and strict schemas are refused; temperature 0 to 1.
OpenRouterCompatible subset, streaming or not. Reasoning text isn't returned (its tokens count as output).
vLLM, SGLang, OpenAI-compatibleCompatible subset; max_completion_tokens is sent upstream as max_tokens; strict tool schemas refused. The generic profile refuses required or named tool choice.
OllamaCompatible subset; explicit tool choice is refused.

Failover

If the model has several routes and its routing policy allows more than one attempt, a request that fails before any output (for example a busy provider) can be tried on another route with the same residency label. Each attempt is charged separately. There is never a failover once a stream has started.

On this page