Download OpenAPI specification:
Public customer API. Private Vultr integration routes that require the platform key are not included.
Create a Message using the Anthropic Messages API format. This endpoint accepts Anthropic-native request payloads and returns Anthropic-native responses. The request is adapted internally to the underlying chat completions engine.
System prompt: Use the top-level system parameter for initial instructions. A later text-only system message may follow a user message; it may be last or precede an assistant message.
Content blocks: Message content may be a string or an array of typed content blocks (text, image, tool_use, tool_result).
Assistant prefill: If the final message uses the assistant role, the response continues immediately from that content.
Streaming: Set stream: true to receive server-sent events with typed event names (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop).
Extended thinking: thinking.type: enabled requires budget_tokens below max_tokens and a model that supports enforced thinking budgets; other models reject this mode. Check reasoning.supports_max_tokens in /v1/models?format=openrouter (in the default list, a max_tokens property under the text output's reasoning parameter). adaptive uses the model's own thinking behavior, and disabled turns thinking off. Thinking blocks use API-generated signatures, and display: omitted hides their text while preserving a signature for continuation.
System instructions: A later system message keeps its position when forwarded to the model. Turn-scoped clear_at and per-message output_config are not supported.
Effort: Top-level output_config.effort maps through the selected model's reasoning_config. The :express model variant derives a thinking budget when the selected model supports enforced budgets and no budget is supplied.
Tools: Tools use input_schema (JSON Schema) rather than parameters. Tool results are sent as tool_result content blocks in subsequent user messages.
Errors: Errors use Anthropic's body, {"type": "error", "error": {"type", "message"}}. A request the gateway refuses is a 400 invalid_request_error; Anthropic has no 422. An unknown model is a 404 not_found_error, a rate-limited request a 429 rate_limit_error with Retry-After, a bad or missing key a 401 authentication_error, and a failure on our side or the engine's a 5xx api_error.
| model required | string The model identifier to use for generation. |
| max_tokens required | integer >= 0 Total completion token allowance, including thinking and visible output. The gateway forwards zero if supplied; whether zero is accepted depends on the upstream engine. |
required | Array of objects (Anthropic Message Param) non-empty Input messages. Models operate on alternating If the final message uses the A text-only |
string or Array of Anthropic Text Block Param (objects) System prompt providing context or instructions to the model. May be a string or an array of text blocks. | |
| stream | boolean Whether to stream the response using server-sent events. Streaming events use typed event names: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop. |
| temperature | number <float> [ 0 .. 1 ] Controls randomness in the model responses. Valid range: 0 to 1. |
| top_p | number <float> [ 0 .. 1 ] Enables nucleus sampling with a cumulative probability cutoff. |
| top_k | integer Limits sampling to the top K options. |
| stop_sequences | Array of strings Custom sequences that, when encountered in the generated text, cause the model to stop. |
object Additional metadata about the request. | |
object Output controls. Effort is translated through the selected model's reasoning configuration. Unsupported configured levels return 400. | |
object Thinking configuration. The adapter signs its own thinking blocks for continuation; it does not emit redacted thinking blocks. | |
Array of objects (Anthropic Tool Definition) Definitions for tools available to the model. | |
object Instructs how the model should use any provided tools. |
{- "model": "string",
- "max_tokens": 0,
- "messages": [
- {
- "role": "user",
- "content": "string",
- "clear_at": "never"
}
], - "system": "string",
- "stream": true,
- "temperature": 0,
- "top_p": 0,
- "top_k": 0,
- "stop_sequences": [
- "string"
], - "metadata": {
- "user_id": "string"
}, - "output_config": {
- "effort": "minimal"
}, - "thinking": {
- "type": "enabled",
- "budget_tokens": 1024,
- "display": "summarized"
}, - "tools": [
- {
- "name": "string",
- "description": "string",
- "input_schema": { },
- "type": "string",
- "cache_control": {
- "type": "ephemeral"
}
}
], - "tool_choice": {
- "type": "auto",
- "name": "string"
}
}{- "id": "string",
- "type": "message",
- "role": "assistant",
- "model": "string",
- "content": [
- {
- "type": "text",
- "text": "string",
- "thinking": "string",
- "signature": "string",
- "id": "string",
- "name": "string",
- "input": { }
}
], - "stop_reason": "end_turn",
- "stop_sequence": "string",
- "usage": {
- "input_tokens": 0,
- "output_tokens": 0,
- "cache_creation_input_tokens": 0,
- "cache_read_input_tokens": 0
}
}Create Chat Completion on a specified text generation model.
Continuation: If the final message has the assistant role, the model continues that message instead of starting a new turn (the same convention as Anthropic and OpenRouter). Send a response that ended with finish_reason length back as the final assistant message to resume it. The message must be non-empty text without tool_calls, trailing whitespace is trimmed, thinking is disabled for the continuation, and only the new text is returned.
Response shape: Responses follow OpenRouter's chat completion shape whichever engine serves the model. Reasoning is always in reasoning (never reasoning_content), usage always has prompt_tokens_details.cached_tokens and completion_tokens_details.reasoning_tokens, ids start with chatcmpl- and tool call ids with call_, finish_reason is one of stop, length, tool_calls, content_filter or error with the engine's own value in native_finish_reason, and engine-specific fields are removed. A -normalize model suffix is accepted for compatibility and ignored.
Errors: Errors use OpenRouter's envelope, {"error": {"code", "message", "type", "metadata": {"error_type"}}}. A streaming request that fails before the first token gets the real HTTP status; one that fails later receives a final chunk carrying error.
Reasoning budget: Set reasoning.max_tokens to limit thinking tokens. This limit must be below max_completion_tokens (or the effective context-clipped limit). Models with reasoning.supports_max_tokens: true in /v1/models?format=openrouter (a max_tokens property under the reasoning parameter in the default list) enforce the budget; other models accept the field without enforcing it. Requested effort maps through the model's reasoning_config when configured. The :express model variant derives a budget for models with enforced budget support when no budget is supplied.
| model | string The model to use. If omitted, the gateway resolves its default model. A |
required | Array of objects (Chat Completion Request Message) non-empty The message context to use for the chat completion request, separated by system, user, assistant, tool, and developer roles. A final |
| continue_final_message | boolean Default: false Continue the final message instead of starting a new assistant turn. Implied whenever the final message has the |
| stream | boolean Indicates whether the response should be streamed. |
| max_tokens | integer >= 1 Deprecated name for |
| max_completion_tokens | integer >= 1 Default: 32768 Maximum completion tokens, including reasoning and visible output. A value below 1 returns 400; omit the field to get the default. A value that does not fit the model's context window with the prompt is not refused. It is cut to the window less the prompt and a 32 token margin, and |
| reasoning_effort | string Enum: "ultra" "max" "xhigh" "high" "medium" "low" "minimal" "none" Shorthand for |
object (Reasoning Request Options) Reasoning controls. Configured models map effort through their database profile. Models advertising | |
| n | integer >= 1 Default: 1 The number of chat completion choices to generate for each input message. |
| seed | integer If you would like a different response from the same message, changing the seed will change the response. A null value generates a random seed. |
| top_k | integer >= 0 Sample only from the k most likely tokens at each step. 0 disables it. Passed to the engine; a model's default is the |
| repetition_penalty | number <float> [ 0 .. 2 ] Penalises tokens already in the prompt or the answer. 1.0 is no penalty and above 1.0 discourages repetition. Passed to the engine, which also reads it as a model default. |
object Constrains the answer's format. | |
| temperature | number <float> [ 0 .. 2 ] Default: 1 A value between 0.0 and 2.0 that controls the randomness of the model's output. When set closer to 1, such as 0.8, the outcome is more unpredictable and creative. Values nearing 0, like 0.2, produce more predictable and less creative results. A temperature of zero requests greedy decoding; exact reproducibility depends on the model engine. |
| top_p | number <float> [ 0 .. 1 ] Default: 1 A value between 0.0 and 1.0 that controls the probability of the model generating a particular token. A higher value will result in more diverse outputs, while a lower value will result in more repetitive outputs. |
| frequency_penalty | number <float> [ -2 .. 2 ] Default: 0 A value between -2.0 and 2.0 that controls how much the model penalizes generating repetitive responses. |
| presence_penalty | number <float> [ -2 .. 2 ] Default: 0 A value between -2.0 and 2.0 that controls how much the model penalizes generating responses that contain certain words or phrases. |
string or Array of strings The strings the model stops generating at, up to 4. A single string is accepted as shorthand for a one-item array. | |
| logprobs | boolean Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. |
| top_logprobs | integer [ 0 .. 20 ] An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with an associated log probability. Optional: |
string or Chat Completion Function Tool Choice (object) or Allowed tools (object) Controls which (if any) tool is called by the model. A tool choice other than | |
Array of objects (Chat Completion Function Tool) A list of function tools the model may call. |
{- "model": "string",
- "messages": [
- {
- "role": "system",
- "content": "string"
}
], - "continue_final_message": false,
- "stream": true,
- "max_tokens": 1,
- "max_completion_tokens": 32768,
- "reasoning_effort": "ultra",
- "reasoning": {
- "effort": "ultra",
- "enabled": true,
- "summary": "auto",
- "max_tokens": 1
}, - "n": 1,
- "seed": 0,
- "top_k": 0,
- "repetition_penalty": 0,
- "response_format": {
- "type": "text",
- "json_schema": {
- "name": "string",
- "strict": true,
- "schema": { }
}
}, - "temperature": 1,
- "top_p": 1,
- "frequency_penalty": 0,
- "presence_penalty": 0,
- "stop": "string",
- "logprobs": true,
- "top_logprobs": 0,
- "tool_choice": "none",
- "tools": [
- {
- "type": "string",
- "function": {
- "name": "string",
- "description": "string",
- "parameters": { },
- "strict": true
}
}
]
}{- "id": "string",
- "object": "string",
- "created": 0,
- "model": "string",
- "choices": [
- {
- "index": 0,
- "message": {
- "role": "system",
- "content": "string",
- "reasoning": "string",
- "tool_calls": [
- {
- "id": "string",
- "type": "string",
- "function": {
- "name": "string",
- "arguments": "string"
}
}
]
}, - "logprobs": {
- "content": [
- {
- "token": "string",
- "logprob": 0.1,
- "bytes": [
- 0
], - "top_logprobs": [
- {
- "token": "string",
- "logprob": 0.1,
- "bytes": [
- 0
]
}
]
}
]
}, - "finish_reason": "stop",
- "native_finish_reason": "string"
}
], - "usage": {
- "completion_tokens": 0,
- "prompt_tokens": 0,
- "total_tokens": 0,
- "prompt_tokens_details": {
- "cached_tokens": 0
}, - "completion_tokens_details": {
- "reasoning_tokens": 0
}
}
}Create a chat completion with context retrieved from items or files in a vector store collection. Reasoning effort, thinking budget, and the express variant follow the Chat Completions behavior.
| collection required | string The vector store collection to search for relevant context. |
| model required | string The model that will be inferred for chat completion. A |
required | Array of objects (Chat Completion Request Message) non-empty The message context to use for the chat completion request, separated by system, user, assistant, tool, and developer roles. |
| max_tokens | integer Deprecated name for |
| max_completion_tokens | integer >= 1 Default: 32768 Maximum completion tokens, including reasoning and visible output. |
| reasoning_effort | string Enum: "ultra" "max" "xhigh" "high" "medium" "low" "minimal" "none" Shorthand for |
object (Reasoning Request Options) Reasoning controls. Configured models map effort through their database profile. Models advertising | |
| n | integer Default: 1 The number of chat completion choices to generate for each input message. |
| seed | integer If you would like a different response from the same message, changing the seed will change the response. A null value generates a random seed. |
| temperature | number <float> [ 0 .. 2 ] Default: 1 A value between 0.0 and 2.0 that controls the randomness of the model's output. When set closer to 1, such as 0.8, the outcome is more unpredictable and creative. Values nearing 0, like 0.2, produce more predictable and less creative results. A temperature of zero requests greedy decoding; exact reproducibility depends on the model engine. |
| top_p | number <float> [ 0 .. 1 ] Default: 1 A value between 0.0 and 1.0 that controls the probability of the model generating a particular token. A higher value will result in more diverse outputs, while a lower value will result in more repetitive outputs. |
string or Array of strings The strings the model stops generating at, up to 4. A single string is accepted as shorthand for a one-item array. | |
| frequency_penalty | number <float> [ -2 .. 2 ] Default: 0 A value between -2.0 and 2.0 that controls how much the model penalizes generating repetitive responses. |
| presence_penalty | number <float> [ -2 .. 2 ] Default: 0 A value between -2.0 and 2.0 that controls how much the model penalizes generating responses that contain certain words or phrases. |
| stream | boolean Indicates whether the response should be streamed. |
| logprobs | boolean Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. |
| top_logprobs | integer [ 0 .. 20 ] An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with an associated log probability. Optional: |
string or Chat Completion Function Tool Choice (object) or Allowed tools (object) Controls which (if any) tool is called by the model. A tool choice other than | |
Array of objects (Chat Completion Function Tool) A list of function tools the model may call. |
{- "collection": "string",
- "model": "string",
- "messages": [
- {
- "role": "system",
- "content": "string"
}
], - "max_tokens": 0,
- "max_completion_tokens": 32768,
- "reasoning_effort": "ultra",
- "reasoning": {
- "effort": "ultra",
- "enabled": true,
- "summary": "auto",
- "max_tokens": 1
}, - "n": 1,
- "seed": 0,
- "temperature": 1,
- "top_p": 1,
- "stop": "string",
- "frequency_penalty": 0,
- "presence_penalty": 0,
- "stream": true,
- "logprobs": true,
- "top_logprobs": 0,
- "tool_choice": "none",
- "tools": [
- {
- "type": "string",
- "function": {
- "name": "string",
- "description": "string",
- "parameters": { },
- "strict": true
}
}
]
}{- "id": "string",
- "object": "string",
- "created": 0,
- "model": "string",
- "choices": [
- {
- "index": 0,
- "message": {
- "role": "system",
- "content": "string",
- "reasoning": "string",
- "tool_calls": [
- {
- "id": "string",
- "type": "string",
- "function": {
- "name": "string",
- "arguments": "string"
}
}
]
}, - "logprobs": {
- "content": [
- {
- "token": "string",
- "logprob": 0.1,
- "bytes": [
- 0
], - "top_logprobs": [
- {
- "token": "string",
- "logprob": 0.1,
- "bytes": [
- 0
]
}
]
}
]
}, - "finish_reason": "stop",
- "native_finish_reason": "string"
}
], - "usage": {
- "completion_tokens": 0,
- "prompt_tokens": 0,
- "total_tokens": 0,
- "prompt_tokens_details": {
- "cached_tokens": 0
}, - "completion_tokens_details": {
- "reasoning_tokens": 0
}
}
}Create an OpenAI-compatible response. reasoning.effort maps through the selected model's reasoning_config when configured. reasoning.max_tokens caps thinking tokens when the model advertises reasoning.supports_max_tokens: true in /v1/models?format=openrouter (a max_tokens property under the reasoning parameter in the default list); other models accept the budget without enforcing it. reasoning.context and reasoning.mode are accepted but ignored. The :express model variant derives a budget for models with enforced budget support when no budget is supplied.
Response shape: Responses follow OpenAI's Responses API object whichever engine serves the model, with the fields OpenRouter's SDKs require always present: completed_at is a unix timestamp once the response has finished, error is null or {code, message}, and text is an object. Function call ids start with call_, and engine-internal fields (kv_transfer_params, input_messages, output_messages, ...) are removed from the response and from every streamed event.
Tools: Only function tools are supported. Any other tool type (web_search, file_search, code_interpreter, mcp, ...) is refused with a 400 naming it.
Stored responses: Nothing is stored. store is accepted and ignored. previous_response_id, conversation and background: true are refused with a 400; send the whole conversation in input. GET and DELETE on /v1/responses/{id} answer a 404 saying so.
Context window: If input plus max_output_tokens passes the model's context length, max_output_tokens is reduced to fit; if the input alone does not fit the request is a 400 with metadata.error_type context_length_exceeded.
Errors: Errors use OpenRouter's envelope, {"error": {"code", "message", "type", "metadata": {"error_type"}}}. A streaming request that fails before the first event gets the real HTTP status and the JSON envelope; one that fails later ends the stream with an error event and a response.failed event.
| model | string If omitted, the gateway resolves its default model. A |
required | string or Array of objects Input text or structured input items. |
| instructions | string A system prompt. |
| max_output_tokens | integer >= 1 Default: 32768 Maximum output tokens, including reasoning and visible output. Reduced to fit the model's context window with the input. |
| temperature | number [ 0 .. 2 ] |
| top_p | number [ 0 .. 1 ] |
Array of objects Function tools only. Any other | |
string or object
| |
object (Reasoning Request Options) Reasoning controls. Configured models map effort through their database profile. Models advertising | |
| store | boolean Accepted and ignored. Nothing is stored, whatever the value. |
| previous_response_id | string or null Refused with a 400 unless null. Responses are not stored; send the whole conversation in |
| conversation | any Refused with a 400. Conversations are not stored. |
| background | boolean
|
| stream | boolean |
{- "model": "string",
- "input": "string",
- "instructions": "string",
- "max_output_tokens": 32768,
- "temperature": 0,
- "top_p": 0,
- "tools": [
- {
- "type": "function",
- "name": "string",
- "description": "string",
- "parameters": { },
- "strict": true
}
], - "tool_choice": "none",
- "reasoning": {
- "effort": "ultra",
- "enabled": true,
- "summary": "auto",
- "max_tokens": 1
}, - "store": true,
- "previous_response_id": "string",
- "conversation": null,
- "background": true,
- "stream": true
}{- "id": "string",
- "object": "response",
- "created_at": 0,
- "completed_at": 0,
- "status": "completed",
- "error": {
- "code": "string",
- "message": "string"
}, - "incomplete_details": {
- "reason": "string"
}, - "model": "string",
- "output": [
- { }
], - "store": true,
- "previous_response_id": "string",
- "text": { },
- "tool_choice": null,
- "tools": [
- { }
], - "usage": {
- "input_tokens": 0,
- "input_tokens_details": {
- "cached_tokens": 0,
- "cache_write_tokens": 0
}, - "output_tokens": 0,
- "output_tokens_details": {
- "reasoning_tokens": 0
}, - "total_tokens": 0
}
}Always a 404. Responses are not stored, so there is nothing to retrieve; the error says so, for SDK calls that would otherwise get a bare "no such endpoint".
| id required | string A response ID, |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}Always a 404. Responses are not stored, so there is nothing to delete.
| id required | string A response ID, |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}Embed text in OpenRouter's embeddings shape, which is OpenAI's with input_type and an id added.
Batches: Up to 2048 inputs per request. The API splits them into batches the model's engine accepts and returns the embeddings in input order.
Queries and documents: Instruction-aware models such as qwen3-embedding-4b retrieve better when queries carry an instruction. Set input_type to search_query on queries and the model's instruction is prepended; documents (search_document, or no input_type) are sent unchanged. A query that already starts with the instruction's marker (e.g. Instruct:) is not prefixed again. The prefix's tokens are billed.
Long inputs: An input longer than the model's maximum input length is truncated to it, not rejected.
| model required | string The embeddings model to use. |
required | string or Array of strings or Array of integers or Array of integers A string, a list of strings, a list of token ids, or a list of token id lists. Up to 2048 inputs. Multimodal inputs are rejected. |
| encoding_format | string Default: "float" Enum: "float" "base64"
|
| dimensions | integer >= 1 Truncate embeddings to this many dimensions, for models trained to support it (e.g. Qwen3-Embedding). |
| input_type | string
|
{- "model": "qwen3-embedding-4b",
- "input": [
- "What is the capital of France?"
], - "input_type": "search_query"
}{- "id": "string",
- "object": "string",
- "data": [
- {
- "object": "string",
- "index": 0,
- "embedding": [
- 0
]
}
], - "model": "string",
- "usage": {
- "prompt_tokens": 0,
- "total_tokens": 0
}
}The live embedding models in OpenRouter's public model shape (data, total_count, links), for OpenRouter's embeddings.listModels(). Each entry has output_modalities: ["embeddings"], the context length and the price per input token in pricing.prompt, the same as GET /v1/models lists for it. Public like GET /v1/models.
| offset | integer >= 0 Default: 0 Models to skip. Offset and limit both omitted is the whole list. |
| limit | integer [ 1 .. 1000 ] Default: 500 Models per page, once offset or limit is set. |
{- "data": [
- {
- "id": "glm-5.3",
- "canonical_slug": "string",
- "hugging_face_id": "string",
- "name": "string",
- "created": 0,
- "description": "string",
- "context_length": 0,
- "architecture": {
- "input_modalities": [
- "text"
], - "output_modalities": [
- "text"
], - "modality": "text+image->text",
- "tokenizer": "Other",
- "instruct_type": "string"
}, - "pricing": {
- "prompt": "0.00000055",
- "completion": "0.00000275",
- "internal_reasoning": "string",
- "input_cache_read": "string"
}, - "top_provider": {
- "context_length": 0,
- "max_completion_tokens": 0,
- "is_moderated": true
}, - "per_request_limits": { },
- "supported_parameters": [
- "temperature",
- "top_p",
- "max_tokens",
- "tools"
], - "default_parameters": {
- "temperature": 0,
- "top_p": 0,
- "top_k": 0,
- "frequency_penalty": 0,
- "presence_penalty": 0,
- "repetition_penalty": 0
}, - "supported_voices": [
- "string"
], - "links": {
- "details": "/v1/model/vultr/glm-5.3"
}, - "reasoning": {
- "mandatory": true,
- "supports_max_tokens": true,
- "supported_efforts": [
- "string"
], - "default_effort": "string",
- "default_enabled": true
}
}
], - "total_count": 0,
- "links": {
- "next": "/v1/models?offset=500&limit=500"
}
}Answer typed questions about a state with a decision ("System One") model, in TypeSafe's /v1/systemone shape, so clients written for TypeSafe work by changing their base URL and key.
A decision model reads the state once per question and generates no text: each answer is a probability distribution over the question's allowed answers.
Question types: noul (yes/no) answers noul, the probability of yes. choice picks one of 1 to 255 options and answers the most likely choice, every option's probabilities and a confidence. score rates on 1 to 10 ordered levels and answers score, the probability-weighted level (it can land between levels), with a legend, probabilities and confidence.
Confidence is 1 when the distribution is on one answer and 0 when it is spread evenly: (k * p_max - 1) / (k - 1) for k answers. noul answers carry none; read noul itself.
Long inputs over the model's maximum input length are rejected with 422, never truncated.
Billing: every question's prompt (the state, the question and its options) is billed as input tokens, so questions on a long state each bill the state. No output tokens are billed.
| model required | string The decision model to use. |
required | string or object or Array of any What to decide about. A string, or structured data (an object or array), which the model reads as JSON. |
required | object Up to 64 questions, by id. Each answer comes back under its question's id. |
{- "model": "mica-v0.1-4b",
- "state": {
- "ticket": "Everything is down and we have a demo at noon."
}, - "questions": {
- "urgent": {
- "type": "noul",
- "instructions": "Does the customer need a reply within the hour?"
}, - "department": {
- "type": "choice",
- "instructions": "Which team should handle this?",
- "criteria": {
- "billing": "invoices, payments, refunds",
- "technical": "bugs, outages, errors"
}
}, - "frustration": {
- "type": "score",
- "instructions": "How frustrated is the customer?",
- "criteria": [
- "Calm",
- "Frustrated",
- "Very angry"
]
}
}
}{- "model": "string",
- "answers": {
- "property1": {
- "type": "noul",
- "noul": 0,
- "choice": "string",
- "score": 0,
- "legend": {
- "property1": "string",
- "property2": "string"
}, - "probabilities": {
- "property1": 0,
- "property2": 0
}, - "confidence": 0
}, - "property2": {
- "type": "noul",
- "noul": 0,
- "choice": "string",
- "score": 0,
- "legend": {
- "property1": "string",
- "property2": "string"
}, - "probabilities": {
- "property1": 0,
- "property2": 0
}, - "confidence": 0
}
}, - "usage": {
- "input_tokens": 0,
- "output_tokens": 0
}
}Rank documents by relevance to a query. Requests and responses follow OpenRouter's rerank schema whichever engine serves the model.
| model required | string The rerank model to use. |
| query required | string non-empty The search query to rank the documents against. |
required | Array of strings or objects non-empty The documents to rank, as plain strings or objects with a |
| top_n | integer [ 1 .. 100 ] Return only the highest scoring |
| return_documents | boolean Default: true Extension to OpenRouter's schema. Set |
{- "model": "bge-reranker-v2-m3",
- "query": "What is the capital of France?",
- "documents": [
- "Berlin is the capital of Germany.",
- "Paris is the capital of France."
], - "top_n": 1
}{- "id": "rerank-3f1c9a0d2b7e4c6a8d5e1f02",
- "model": "bge-reranker-v2-m3",
- "results": [
- {
- "index": 1,
- "relevance_score": 0.9998499,
- "document": {
- "text": "Paris is the capital of France."
}
}
], - "usage": {
- "total_tokens": 36
}
}Generate speech from text. OpenAI's request shape plus OpenRouter's, so both SDKs work unchanged. The response body is the audio itself.
Voices: voice selects the mode as well as the voice. A built-in speaker (GET /v1/audio/voices?model= lists them) speaks the input, and instructions can direct its style. design creates a new voice from the description in instructions. clone (or leaving voice out) with input_references speaks in the voice of a reference clip: send it base64-encoded in input_audio.data, with its transcript in a text part for the closest match. Clip URLs are not accepted. A cloned voice cannot follow instructions.
Formats: response_format is mp3 (the default), opus, flac, wav or pcm (24 kHz 16-bit mono, labelled audio/pcm;rate=24000;channels=1). aac returns 400.
Streaming: stream_format audio (the default) returns the audio; pcm at speed 1 streams as it is generated. sse returns speech.audio.delta events with base64 audio and a final speech.audio.done event with usage; pcm at speed 1 streams, other formats arrive as one delta. A failure mid-stream ends an SSE stream with a speech.audio.error event.
Billing: Per second of audio generated, before speed, rounded up. The seconds and cost are in the X-Usage-Seconds and X-Usage-Cost headers, or in the speech.audio.done usage; a raw pcm stream reports neither.
OpenRouter's provider, user, session_id and trace are accepted and ignored, except that provider.options under our OpenRouter slug can carry instructions, language, speed, seed, temperature, top_p and top_k.
| model required | string The text-to-speech model, e.g. |
| input required | string The text to speak, up to the model's |
string or object A built-in speaker, | |
| instructions | string Style for a built-in speaker, or the description of the voice for |
Array of objects OpenRouter's reference parts for | |
| language | string The language of the input, as |
| response_format | string Default: "mp3" Enum: "mp3" "opus" "flac" "wav" "pcm" |
| stream_format | string Default: "audio" Enum: "audio" "sse" |
| speed | number [ 0.25 .. 4 ] Default: 1 |
| seed | integer |
| temperature | number [ 0 .. 2 ] |
| top_p | number [ 0 .. 1 ] The engine takes values above 0, so 0 is sent as 0.000001, which keeps only the likeliest token. |
| top_k | integer >= -1 |
{- "model": "qwen3-tts-12hz-1.7b",
- "input": "Welcome to Vultr Serverless Inference.",
- "voice": "aiden",
- "instructions": "Warm and upbeat, like a radio host."
}Transcribe audio. The request can be OpenRouter's JSON form, with the audio base64-encoded in input_audio, or OpenAI's multipart/form-data upload with the audio in file. Both return the same response, so the OpenAI and OpenRouter SDKs work unchanged.
Audio: Up to 25 MB after decoding and 4 hours long. Audio URLs are not accepted.
Formats: response_format is json (the default) or text, which returns the bare transcript as text/plain. verbose_json, timestamp_granularities, srt and vtt are not supported and return 400.
Streaming: With stream: true the response is OpenAI's server-sent events: transcript.text.delta events with each piece of the transcript as it is decoded, then one transcript.text.done with the whole text and the same usage as the JSON response. A failure after the stream began ends it with a transcript.text.error event in the error envelope. The upload still completes before decoding starts; this is not live audio.
Billing: Per second of audio, rounded up, as reported in usage.seconds.
| model required | string The speech-to-text model to use, e.g. |
required | object |
| language | string ISO-639-1 code of the spoken language. Detected when omitted. |
| prompt | string Text to guide spelling and style, such as names and terms in the audio. |
| temperature | number [ 0 .. 1 ] |
| response_format | string Default: "json" Enum: "json" "text" |
| stream | boolean Default: false Stream the transcript as OpenAI's |
{- "model": "whisper-large-v3-turbo",
- "input_audio": {
- "data": "UklGRiQAAABXQVZFZm10IBAAAAABAAEAgD4AAAB9AAACABAAZGF0YQAAAAA=",
- "format": "wav"
}, - "language": "en"
}{- "text": "string",
- "usage": {
- "type": "string",
- "seconds": 0
}
}A speech model's voices: its built-in speakers, and the design and clone voices that select voice design and cloning.
| model required | string The text-to-speech model, e.g. |
{- "object": "list",
- "model": "string",
- "data": [
- {
- "id": "string",
- "name": "string",
- "mode": "string",
- "native_language": "string",
- "description": "string"
}
]
}Creates a vector store collection for searchable embeddings.
| name required | string The name of the vector store collection. This is also used to auto-generate a unique ID for the record. |
{- "name": "string"
}{- "collection": {
- "id": "string",
- "name": "string",
- "created": "string"
}
}Updates a vector store collection record.
| id required | string The ID of the vector store collection. |
| name required | string The name of the vector store collection. Note: the previously generated unique ID will remain the same. |
{- "name": "string"
}{- "collection": {
- "id": "string",
- "name": "string",
- "created": "string"
}
}Deletes a vector store collection record. This will also remove all items in the collection.
| id required | string The ID of the vector store collection. |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}Searches items in a vector store collection for the closest embeddings matches.
| id required | string The ID of the vector store collection. |
| input required | string The text query to search against the embeddings items in the vector store collection. |
{- "input": "string"
}{- "results": [
- {
- "id": "string",
- "created": "string",
- "content": "string"
}
], - "usage": {
- "prompt_tokens": 0,
- "total_tokens": 0
}
}Retrieve a list of items within a vector store collections.
| id required | string The ID of the vector store collection. |
{- "items": [
- {
- "id": "string",
- "created": "string",
- "description": "string"
}
]
}Adds an item to a vector store collection.
| id required | string The ID of the vector store collection. |
| content required | string The text to be converted into embeddings and stored in the vector store collection. |
| description | string A description of the contents in this collection item record. If omitted, this value will default to a shortened version of the text stored in the collection. |
| auto_chunk | boolean Indicates whether the system will automatically chunk the content if it exceeds the embeddings model's maximum sequence length. If set to true, the content will be split into 300 token chunks with 20 tokens of overlap before and after each piece. |
{- "content": "string",
- "description": "string",
- "auto_chunk": true
}{- "item": {
- "id": "string",
- "created": "string",
- "description": "string",
- "content": "string"
}, - "usage": {
- "prompt_tokens": 0,
- "total_tokens": 0
}
}Retrieve a vector store collection item by the ID.
| id required | string The ID of the vector store collection. |
| itemid required | string The ID of the vector store collection item. |
{- "item": {
- "id": "string",
- "created": "string",
- "description": "string",
- "content": "string"
}
}Updates a vector store collection item record.
| id required | string The ID of the vector store collection. |
| itemid required | string The ID of the vector store collection item. |
| description required | string A description of the contents in this collection item record. |
{- "description": "string"
}{- "item": {
- "id": "string",
- "created": "string",
- "description": "string",
- "content": "string"
}
}Deletes a vector store collection item record.
| id required | string The ID of the vector store collection. |
| itemid required | string The ID of the vector store collection item. |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}Retrieve a list of files within a vector store collections.
| id required | string The ID of the vector store collection. |
{- "files": [
- {
- "id": "string",
- "filename": "string",
- "status": "enqueued",
- "error": "string",
- "items": 0,
- "tokens": 0
}
]
}Adds a file to a vector store collection.
| id required | string The ID of the vector store collection. |
| file required | string <binary> The file object to be uploaded to the vector store collection. |
{- "file": {
- "id": "string",
- "filename": "string",
- "status": "enqueued",
- "error": "string",
- "items": 0,
- "tokens": 0
}
}Retrieve a vector store collection file by the ID.
| id required | string The ID of the vector store collection. |
| fileid required | string The ID of the vector store collection file. |
{- "file": {
- "id": "string",
- "filename": "string",
- "status": "enqueued",
- "error": "string",
- "items": 0,
- "tokens": 0
}
}Deletes a vector store collection file record.
| id required | string The ID of the vector store collection. |
| fileid required | string The ID of the vector store collection file. |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}The image models currently served, in OpenRouter's image model list shape, for images.listModels() in OpenRouter's SDKs. Each entry names the parameters the model accepts (supported_parameters: enum, range and boolean descriptors) and links its endpoint record in endpoints. Everything comes from the model's capability spec, so the limits listed are the limits enforced.
size is a constraint, not a list: a WxH string whose sides are multiples of multiple_of, whose area is between min_area and max_area pixels, and whose longer side is at most max_ratio times the shorter. aspect_ratio and resolution are enums, and resolution.sizes is the table of the size each aspect ratio and resolution gives. Prices are on GET /v1/models (per megapixel).
{- "data": [
- {
- "id": "z-image-turbo",
- "name": "Z-Image Turbo",
- "description": "string",
- "created": 0,
- "endpoints": "/v1/images/models/vultr/z-image-turbo/endpoints",
- "architecture": {
- "input_modalities": [
- "text"
], - "output_modalities": [
- "image"
]
}, - "supported_parameters": {
- "property1": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}, - "property2": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}
}, - "supports_streaming": true
}
]
}The endpoint record of one image model, the endpoints link of its list entry. author is not checked; the gateway's model ids have none, and vultr is the documented one. allowed_passthrough_parameters are the names the model takes under provider.options (vultr or any other slug, flat or under parameters; OpenRouter's SDKs keep only stealth): every declared parameter and alias, steps included. They are also accepted at the top level of the request.
| author required | string Example: vultr |
| slug required | string Example: z-image-turbo |
{- "id": "string",
- "endpoints": [
- {
- "provider_name": "Vultr",
- "provider_slug": "vultr",
- "provider_tag": "string",
- "supported_parameters": {
- "property1": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}, - "property2": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}
}, - "allowed_passthrough_parameters": [
- "string"
], - "passthrough_parameters": {
- "property1": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}, - "property2": {
- "type": "enum",
- "values": [
- "string"
], - "min": 0,
- "max": 0,
- "default": null
}
}, - "pricing": [
- {
- "billable": "output_image",
- "cost_usd": 0,
- "unit": "megapixel"
}
], - "supports_streaming": true
}
]
}OpenRouter's image endpoint, for images.generate() in OpenRouter's SDKs (set the base URL to this gateway's /v1). It is the same implementation as /v1/images/generations and /v1/images/edits; what differs is the default answer: base64 (response_format b64_json), because OpenRouter's SDKs require b64_json in every item. Send input_references to edit with a model that has reference editing.
Model-specific settings (steps, guidance_scale, negative_prompt) go at the top level, or under provider.options for an SDK that cannot add fields: OpenRouter's SDKs keep only the slugs on their closed list, so use provider.options.stealth there. A key the model does not declare is a 400 naming it.
Billing is output megapixels only (usage.megapixels, at the model's price per megapixel in usage.cost); reference images are not billed, and a failed request is not billed.
| prompt required | string <= 2000 characters A text description of the desired image(s). The maximum length is the model's (2000 characters for |
| model | string Default: "z-image-turbo" The model to use for image generation. Defaults to |
| n | integer [ 1 .. 10 ] Default: 1 The number of images to generate. The images in one request are generated together, so changing |
| response_format | string Enum: "url" "b64_json" How the image comes back. |
| size | string Default: "1024x1024"
|
| width | integer Deprecated Legacy spelling of |
| height | integer Deprecated See |
| aspect_ratio | string One of the model's aspect ratios ( |
| resolution | string One of the model's resolutions ( |
| seed | integer <int64> [ -9223372036854776000 .. 9223372036854776000 ] Random seed. Must be an integer; fractional values are rejected rather than rounded. When omitted, each request uses a random seed. Repeating a request with the same seed and the same value for every other field, including |
| output_format | string Default: "png" Enum: "png" "jpeg" "webp" |
| output_compression | integer [ 0 .. 100 ] The encoder quality for |
| background | string Default: "auto" Enum: "auto" "opaque" "transparent"
|
Array of objects OpenRouter's reference images, for a model with reference editing ( | |
| steps | integer [ 1 .. 50 ] The number of denoising steps ( For |
| num_inference_steps | integer [ 1 .. 50 ] Alias of |
| guidance_scale | number <float> [ 0 .. 20 ] Classifier-free guidance strength. When omitted, the model's default is used: 0 for
|
| negative_prompt | string Text describing what to steer the image away from. It only takes effect when |
| provider | object OpenRouter's provider preferences. The routing fields ( |
| stream | boolean
|
| partial_images | integer [ 0 .. 0 ] OpenAI's partial image count. Only 0 is accepted; a stream carries the finished image only. |
| user | string Accepted and ignored. |
| session_id | string Accepted and ignored. |
| trace | object Accepted and ignored. |
{- "model": "z-image-turbo",
- "prompt": "Studio photograph of a red ceramic teapot beside a blue cup on a wooden table.",
- "aspect_ratio": "16:9",
- "resolution": "1K",
- "provider": {
- "options": {
- "stealth": {
- "steps": 8
}
}
}
}{- "created": 0,
- "data": [
- {
- "b64_json": "string",
- "url": "string",
- "media_type": "image/png",
- "revised_prompt": "string"
}
], - "size": "1024x1024",
- "output_format": "png",
- "background": "opaque",
- "usage": {
- "prompt_tokens": 0,
- "completion_tokens": 0,
- "total_tokens": 0,
- "input_tokens": 0,
- "output_tokens": 0,
- "input_tokens_details": {
- "text_tokens": 0,
- "image_tokens": 0
}, - "cost": 0.02097152,
- "megapixels": 1.048576
}
}Creates one or more images from a prompt, in OpenAI's shape (images.generate() in OpenAI's SDKs). The default model is z-image-turbo. GET /v1/images/models lists the image models currently served with what each accepts.
The answer is a link (url) unless response_format is b64_json; POST /v1/images answers base64 by default. Every error is OpenRouter's envelope.
quality, moderation and style are not supported and are a 400 naming the field. stream: true answers one image_generation.completed event and [DONE] (partial_images above 0 is a 400). output_format (png, jpeg, webp) and output_compression are passed to the engine. Model-specific settings are steps (or num_inference_steps), guidance_scale and negative_prompt, also under provider.options (see POST /v1/images).
For reproducible output, send an integer seed together with steps and guidance_scale, and keep every other field the same between requests.
| prompt required | string <= 2000 characters A text description of the desired image(s). The maximum length is the model's (2000 characters for |
| model | string Default: "z-image-turbo" The model to use for image generation. Defaults to |
| n | integer [ 1 .. 10 ] Default: 1 The number of images to generate. The images in one request are generated together, so changing |
| response_format | string Enum: "url" "b64_json" How the image comes back. |
| size | string Default: "1024x1024"
|
| width | integer Deprecated Legacy spelling of |
| height | integer Deprecated See |
| aspect_ratio | string One of the model's aspect ratios ( |
| resolution | string One of the model's resolutions ( |
| seed | integer <int64> [ -9223372036854776000 .. 9223372036854776000 ] Random seed. Must be an integer; fractional values are rejected rather than rounded. When omitted, each request uses a random seed. Repeating a request with the same seed and the same value for every other field, including |
| output_format | string Default: "png" Enum: "png" "jpeg" "webp" |
| output_compression | integer [ 0 .. 100 ] The encoder quality for |
| background | string Default: "auto" Enum: "auto" "opaque" "transparent"
|
Array of objects OpenRouter's reference images, for a model with reference editing ( | |
| steps | integer [ 1 .. 50 ] The number of denoising steps ( For |
| num_inference_steps | integer [ 1 .. 50 ] Alias of |
| guidance_scale | number <float> [ 0 .. 20 ] Classifier-free guidance strength. When omitted, the model's default is used: 0 for
|
| negative_prompt | string Text describing what to steer the image away from. It only takes effect when |
| provider | object OpenRouter's provider preferences. The routing fields ( |
| stream | boolean
|
| partial_images | integer [ 0 .. 0 ] OpenAI's partial image count. Only 0 is accepted; a stream carries the finished image only. |
| user | string Accepted and ignored. |
| session_id | string Accepted and ignored. |
| trace | object Accepted and ignored. |
{- "model": "z-image-turbo",
- "prompt": "Studio photograph of a red ceramic teapot beside a blue cup on a wooden table.",
- "n": 1,
- "size": "1024x1024",
- "response_format": "b64_json"
}{- "created": 0,
- "data": [
- {
- "b64_json": "string",
- "url": "string",
- "media_type": "image/png",
- "revised_prompt": "string"
}
], - "size": "1024x1024",
- "output_format": "png",
- "background": "opaque",
- "usage": {
- "prompt_tokens": 0,
- "completion_tokens": 0,
- "total_tokens": 0,
- "input_tokens": 0,
- "output_tokens": 0,
- "input_tokens_details": {
- "text_tokens": 0,
- "image_tokens": 0
}, - "cost": 0.02097152,
- "megapixels": 1.048576
}
}Edits or recombines images with a prompt, in OpenAI's shape (images.edit() in OpenAI's SDKs): multipart image[] files and an optional mask, or JSON images: [{image_url}] and mask: {image_url}; OpenRouter's input_references works too. The same implementation as /v1/images/generations, and the same fields and answer.
Only a model with reference editing takes this route. z-image-turbo has none: the request is a 400 naming image (or images, input_references, mask, whichever was sent). Each image is an upload, a base64 data URI or an https URL the gateway fetches itself (public addresses only); the engine only receives bytes. An image's type is read from its bytes and checked against the model's accepted formats.
Billing is the output megapixels only, so an edit with any number of references bills the same as a generation of the same size. size auto is the model's default size, not the first image's.
| prompt required | string |
| model | string |
| image required | Array of strings <binary> [ items <binary > ] One or more image files. Send them as |
| mask | string <binary> A PNG mask; it needs at least one image. |
| n | integer |
| size | string |
| response_format | string Enum: "url" "b64_json" |
| output_format | string |
| output_compression | integer |
| background | string |
| seed | integer |
| stream | boolean |
{- "created": 0,
- "data": [
- {
- "b64_json": "string",
- "url": "string",
- "media_type": "image/png",
- "revised_prompt": "string"
}
], - "size": "1024x1024",
- "output_format": "png",
- "background": "opaque",
- "usage": {
- "prompt_tokens": 0,
- "completion_tokens": 0,
- "total_tokens": 0,
- "input_tokens": 0,
- "output_tokens": 0,
- "input_tokens_details": {
- "text_tokens": 0,
- "image_tokens": 0
}, - "cost": 0.02097152,
- "megapixels": 1.048576
}
}Starts a video render and answers 202 at once with the job; poll its polling_url, or pass callback_url to be notified. The request follows OpenRouter's video API, so the OpenRouter SDKs work by changing the base URL and key. OpenAI/Sora field names are accepted too: seconds for duration, size, and input_reference (an uploaded file in multipart/form-data, an https URL string, or {image_url}) for the first frame. GET /v1/videos/models lists the models with what each accepts.
Modes (minimax-h3): text only; frame_images with a first_frame, a last_frame or both; or input_references (up to 9 images, 3 videos and 3 audio clips, 12 in all). When both frame_images and input_references are sent, frame_images is used.
Media: an https URL or a base64 data URI. Images JPG, PNG, WEBP, HEIC or HEIF up to 30 MB; videos MP4 or MOV up to 50 MB, 2-15 s; audio WAV or MP3 up to 15 MB, 2-15 s.
Sizes (minimax-h3): 4 to 15 seconds at 24 fps, with stereo AAC audio. Send aspect_ratio or size:
aspect_ratio |
size |
|---|---|
16:9 (default) |
1344x768 |
9:16 |
768x1344 |
4:3 |
1024x768 |
3:4 |
768x1024 |
1:1 |
768x768 |
21:9 |
1536x672 |
Output: duration snaps to the model's frame grid, so a 4 s request is about 4.4 s; read output.duration. Videos have a generated soundtrack unless generate_audio is false.
Prompts are used as written. MiniMax's and fal's APIs rewrite prompts by default, so the same prompt can give a different video here.
Expected wait: while a job is pending, queue_position counts the jobs ahead of it for the same model. Once it starts, a minimax-h3 render takes about 3 seconds per second of video for text and frame jobs, and about 5 seconds per second with input_references: a 6 s text job is ready in about 20 s. progress is estimated from these times.
Steps (minimax-h3): by default a render uses H3's Turbo LoRA: 4 denoising steps, or 8 with input_references. To run the full base model instead, send provider.options.vultr.parameters.num_inference_steps from 10 to 50; more steps take longer and cost more. Sending the Turbo count itself changes nothing.
Billing: per second of requested duration at the model's price (pricing_skus in GET /v1/videos/models), only when the job completes. Steps above the Turbo count scale the billed seconds by steps / Turbo count, rounded up to a whole second: a 6 s text job at 20 steps bills 30 s (6 x 20 / 4). usage.billed_seconds shows the result. Failed and cancelled jobs are not billed.
| model required | string |
| prompt required | string <= 7000 characters Must be a string (a list or number is a 400). |
| duration | integer Seconds of video. minimax-h3 takes 4-15; default 6. A whole number sent as text ( |
| seconds | integer OpenAI/Sora's name for |
| resolution | string |
| aspect_ratio | string Default: "16:9" Enum: "16:9" "9:16" "4:3" "3:4" "1:1" "21:9" |
| size | string Enum: "1344x768" "768x1344" "1024x768" "768x1024" "768x768" "1536x672" Width x height; an alternative to |
Array of objects <= 2 items | |
Array of objects Reference media for the subject, style or sound. Ignored when | |
string or object OpenAI/Sora's first frame, as an https URL (or base64 data URI) string or as | |
| generate_audio | boolean Default: true Strictly |
| seed | integer |
| callback_url | string An https URL notified on completion, failure, cancellation and expiry. See |
| provider | object Model-specific options ( |
{- "model": "minimax-h3",
- "prompt": "A red fox trots through fresh snow at dawn, breath steaming.",
- "duration": 6,
- "aspect_ratio": "16:9",
}{- "id": "video_3f2a9c1e7b4d5a6c8e0f1a2b",
- "object": "video",
- "model": "string",
- "status": "pending",
- "progress": 0,
- "created_at": 0,
- "completed_at": 0,
- "expires_at": 0,
- "polling_url": "string",
- "seconds": "string",
- "size": "string",
- "queue_position": 0,
- "unsigned_urls": [
- "string"
], - "output": {
- "duration": 0,
- "fps": 0,
- "width": 0,
- "height": 0,
- "has_audio": true,
- "count": 0
}, - "usage": {
- "cost": 0,
- "output_seconds": 0,
- "billed_seconds": 0,
- "input_image_count": 0,
- "input_video_seconds": 0,
- "input_audio_seconds": 0
}, - "error": {
- "code": "invalid_request",
- "message": "string"
}
}Your video jobs, newest first.
| limit | integer [ 1 .. 100 ] Default: 20 |
| after | string A job id from the previous page ( |
{- "object": "list",
- "data": [
- {
- "id": "video_3f2a9c1e7b4d5a6c8e0f1a2b",
- "object": "video",
- "model": "string",
- "status": "pending",
- "progress": 0,
- "created_at": 0,
- "completed_at": 0,
- "expires_at": 0,
- "polling_url": "string",
- "seconds": "string",
- "size": "string",
- "queue_position": 0,
- "unsigned_urls": [
- "string"
], - "output": {
- "duration": 0,
- "fps": 0,
- "width": 0,
- "height": 0,
- "has_audio": true,
- "count": 0
}, - "usage": {
- "cost": 0,
- "output_seconds": 0,
- "billed_seconds": 0,
- "input_image_count": 0,
- "input_video_seconds": 0,
- "input_audio_seconds": 0
}, - "error": {
- "code": "invalid_request",
- "message": "string"
}
}
], - "first_id": "string",
- "last_id": "string",
- "has_more": true
}The video models currently served, in OpenRouter's VideoModel shape (a data list, no pagination; OpenRouter's SDKs parse it) with the durations, sizes, inputs and passthrough parameters each accepts, plus our own fields (inputs, operations, supported_fps, frame_grid). Requests are checked against these entries before they are queued. A model without a valid capability spec is not listed.
{- "data": [
- {
- "id": "string",
- "canonical_slug": "string",
- "name": "string",
- "description": "string",
- "hugging_face_id": "string",
- "created": 0,
- "supported_durations": [
- 0
], - "supported_resolutions": [
- "string"
], - "supported_aspect_ratios": [
- "string"
], - "supported_sizes": [
- "string"
], - "supported_frame_images": [
- "first_frame"
], - "generate_audio": true,
- "generates_audio": true,
- "seed": true,
- "upscale_factor": { },
- "creativity": [
- 0
], - "allowed_passthrough_parameters": [
- "string"
], - "pricing_skus": {
- "property1": "string",
- "property2": "string"
}, - "operations": [
- "generate"
], - "inputs": { },
- "supported_fps": [
- 0
], - "frame_grid": "string"
}
]
}The job's status, estimated progress while it renders, and once completed its unsigned_urls, output and usage.
| id required | string The video job's ID, |
{- "id": "video_3f2a9c1e7b4d5a6c8e0f1a2b",
- "object": "video",
- "model": "string",
- "status": "pending",
- "progress": 0,
- "created_at": 0,
- "completed_at": 0,
- "expires_at": 0,
- "polling_url": "string",
- "seconds": "string",
- "size": "string",
- "queue_position": 0,
- "unsigned_urls": [
- "string"
], - "output": {
- "duration": 0,
- "fps": 0,
- "width": 0,
- "height": 0,
- "has_audio": true,
- "count": 0
}, - "usage": {
- "cost": 0,
- "output_seconds": 0,
- "billed_seconds": 0,
- "input_image_count": 0,
- "input_video_seconds": 0,
- "input_audio_seconds": 0
}, - "error": {
- "code": "invalid_request",
- "message": "string"
}
}A pending or in_progress job is cancelled and not billed; it answers with the job, now cancelled. A finished job is deleted with its video and answers {id, object, deleted}.
Cancelling stops billing at once. A render that has already started stops within about 10 seconds.
| id required | string The video job's ID, |
{- "id": "video_3f2a9c1e7b4d5a6c8e0f1a2b",
- "object": "video",
- "model": "string",
- "status": "pending",
- "progress": 0,
- "created_at": 0,
- "completed_at": 0,
- "expires_at": 0,
- "polling_url": "string",
- "seconds": "string",
- "size": "string",
- "queue_position": 0,
- "unsigned_urls": [
- "string"
], - "output": {
- "duration": 0,
- "fps": 0,
- "width": 0,
- "height": 0,
- "has_audio": true,
- "count": 0
}, - "usage": {
- "cost": 0,
- "output_seconds": 0,
- "billed_seconds": 0,
- "input_image_count": 0,
- "input_video_seconds": 0,
- "input_audio_seconds": 0
}, - "error": {
- "code": "invalid_request",
- "message": "string"
}
}The MP4. Available once the job is completed and until expires_at (7 days).
| id required | string The video job's ID, |
| index | integer >= 0 Default: 0 |
{- "error": {
- "code": 0,
- "message": "string",
- "type": "string",
- "metadata": {
- "error_type": "string"
}
}
}The secret your webhooks are signed with, created on first request. It is separate from your API key, so a webhook receiver never holds a key that can spend.
Each delivery is a POST of {type, created_at, data} with data the job as GET /v1/videos/{id} returns it. Events: video.generation.completed, video.generation.failed, video.generation.cancelled, video.generation.expired. Headers:
X-Vultr-Signature: t=<unix>,v1=<hex>, the HMAC-SHA256 of <t>.<raw body> with this secret. Reject old timestamps. During a rotation a second v1= is signed with the previous secret.X-Vultr-Idempotency-Key: <job id>-<status>; retried deliveries repeat it.A 2xx within 10 seconds is delivered; anything else is retried after 30 s, 2 min, 10 min, 1 h and 6 h. Deliveries can arrive more than once or out of order: use the idempotency key, and the job's status, rather than arrival order.
Verifying a delivery. Compute the HMAC over the raw request body, before any JSON parsing, with the whole secret including its whsec_ prefix. Accept the delivery when any v1= value matches.
import hashlib, hmac, time
def verify(raw_body: bytes, header: str, secret: str, tolerance: int = 300) -> bool:
pairs = [part.split("=", 1) for part in header.split(",")]
t = next((v for k, v in pairs if k == "t"), None)
if t is None or abs(time.time() - int(t)) > tolerance:
return False
expected = hmac.new(secret.encode(), t.encode() + b"." + raw_body, hashlib.sha256).hexdigest()
return any(k == "v1" and hmac.compare_digest(expected, v) for k, v in pairs)
const crypto = require("crypto");
function verify(rawBody, header, secret, tolerance = 300) {
const pairs = header.split(",").map((part) => part.split("="));
const t = pairs.find(([k]) => k === "t")?.[1];
if (!t || Math.abs(Date.now() / 1000 - Number(t)) > tolerance) return false;
const expected = crypto.createHmac("sha256", secret).update(`${t}.`).update(rawBody).digest();
return pairs.some(([k, v]) => k === "v1" && v.length === 64 &&
crypto.timingSafeEqual(expected, Buffer.from(v, "hex")));
}
{- "object": "webhook_secret",
- "secret": "whsec_9a1c...",
- "created_at": 0,
- "rotated_at": 0,
- "previous_secret_valid_until": 0
}Replaces the secret. The previous one keeps signing beside it for 24 hours (previous_secret_valid_until).
{- "object": "webhook_secret",
- "secret": "whsec_9a1c...",
- "created_at": 0,
- "rotated_at": 0,
- "previous_secret_valid_until": 0
}Retrieve the public inference model catalog. By default this returns OpenRouter ModelDocumentV2 entries (the provider document); OpenRouter's own SDKs, recognised by their User-Agent, and requests with format=openrouter get OpenRouter's public model list (data, total_count, links) instead, which they can parse. Send the Anthropic-Version header for an Anthropic-compatible model list with cursor pagination. Anthropic mode includes vultr/claude-code/<model-id> aliases because Claude Code only discovers IDs containing claude or anthropic; the original IDs remain available. X-Vultr-Model-Type applies only to the default catalog. TypeSafe's SDKs, which send X-TypeSafe-SDK, get TypeSafe's model list instead: the decision models for POST /v1/systemone as {"models": [{name, description, release_date}]}.
| format | string Value: "openrouter" Set to |
| output_modalities | string Example: output_modalities=speech,transcription Comma-separated output modalities to list: |
| input_modalities | string Example: input_modalities=image Comma-separated input modalities: |
| supported_parameters | string Example: supported_parameters=tools,top_k Comma-separated parameter names, as in the public model list's |
| offset | integer >= 0 OpenRouter public list only. Models to skip. Offset and limit both omitted is the whole list. |
| after_id | string Anthropic mode only. Return the page after this model ID. |
| before_id | string Anthropic mode only. Return the page before this model ID. |
| limit | integer [ 1 .. 1000 ] Default: 20 Anthropic mode: models per page, default 20. OpenRouter public list: models per page, default 500 once offset or limit is set. |
| anthropic-version | string Example: 2023-06-01 When present, return Anthropic model entries and pagination fields instead of the default catalog. |
| X-Vultr-Model-Type | string Value: "chat" Filter the default catalog to chat-completion models. Ignored in Anthropic mode. |
| X-TypeSafe-SDK | string Example: typesafe-sdk/0.7.2 Sent by TypeSafe's SDKs (for example |
{- "data": [
- {
- "schema_version": "2.4",
- "id": "string",
- "name": "string",
- "created": 0,
- "description": "string",
- "datacenters": [
- {
- "country_code": "string",
- "region": "string"
}
], - "hugging_face_id": "string",
- "quantization": "string",
- "input_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "output_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "reasoning": {
- "mandatory": true,
- "default_effort": "string",
- "default_enabled": true,
- "supported_efforts": [
- "ultra"
], - "supports_max_tokens": true
}
}
]
}The number of models the list holds under the same filters, as {"data": {"count": n}}, for OpenRouter's models.count(). Public like the list. The default follows the list: text for OpenRouter's SDKs and format=openrouter, every model otherwise.
| output_modalities | string Example: output_modalities=speech,transcription Comma-separated output modalities, or |
| input_modalities | string Comma-separated input modalities, as on |
| supported_parameters | string Comma-separated parameter names, as on |
| format | string Value: "openrouter" Set to |
{- "data": {
- "count": 12
}
}Retrieves one model: the same entry GET /v1/models lists for it, returned on its own as OpenAI's retrieve does. With an anthropic-version header the answer is Anthropic's model object, for the chat models Anthropic's list shows (including the vultr/claude-code/ aliases). A model ID with a slash can be sent as path segments (/v1/models/zai-org/GLM-5) or percent-encoded (/v1/models/zai-org%2FGLM-5). IDs are case-sensitive. An unknown ID, or a model that is not currently served, is a 404 in the error envelope.
| id required | string The ID of the inference model. |
{- "schema_version": "2.4",
- "id": "string",
- "name": "string",
- "created": 0,
- "description": "string",
- "datacenters": [
- {
- "country_code": "string",
- "region": "string"
}
], - "hugging_face_id": "string",
- "quantization": "string",
- "input_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "output_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "reasoning": {
- "mandatory": true,
- "default_effort": "string",
- "default_enabled": true,
- "supported_efforts": [
- "ultra"
], - "supports_max_tokens": true
}
}Retrieve a model whose identifier contains one slash, using the two path segments as its ID.
| id required | string The ID of the inference model. |
| id2 required | string The second segment of the model identifier. |
{- "schema_version": "2.4",
- "id": "string",
- "name": "string",
- "created": 0,
- "description": "string",
- "datacenters": [
- {
- "country_code": "string",
- "region": "string"
}
], - "hugging_face_id": "string",
- "quantization": "string",
- "input_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "output_modalities": [
- {
- "type": "text",
- "supported_inputs": { },
- "supported_parameters": { },
- "streaming": true,
- "max_length": {
- "value": 0,
- "unit": "string"
}, - "pricing": [
- {
- "type": "string",
- "unit": "string",
- "cost_usd": "string"
}
]
}
], - "reasoning": {
- "mandatory": true,
- "default_effort": "string",
- "default_enabled": true,
- "supported_efforts": [
- "ultra"
], - "supports_max_tokens": true
}
}{- "chat": [
- {
- "id": "string",
- "created": 0,
- "object": "string",
- "owned_by": "string",
- "features": [
- "string"
]
}
], - "audio": [
- {
- "id": "string",
- "created": "string",
- "price": 0.1
}
], - "image": [
- {
- "id": "string",
- "created": 0,
- "price": 0.1
}
]
}View usage information for the current and previous months.
{- "usage": {
- "current_month": {
- "chat_usage": [
- {
- "model": "string",
- "output_tokens": 0,
- "input_tokens": 0
}
], - "tts": 0,
- "tts_sm": 0,
- "image": 0.1,
- "image_sm": 0.1,
- "chat": 0,
- "chat_input": 0
}, - "previous_month": {
- "chat_usage": [
- {
- "model": "string",
- "output_tokens": 0,
- "input_tokens": 0
}
], - "tts": 0,
- "tts_sm": 0,
- "image": 0.1,
- "image_sm": 0.1,
- "chat": 0,
- "chat_input": 0
}
}
}