Skip to main content
Sail provides inference endpoints compatible with the OpenAI Responses API, the OpenAI Chat Completions API, and the Anthropic Messages API. All three inference APIs accept the same models and completion windows. Sail also offers a Batch API for running large numbers of Responses API requests efficiently in one asynchronous job.

Responses API

RecommendedOpenAI SDK compatibleAPI reference

Supported features

Prompt cache routing

For Responses and Chat Completions, reuse the same prompt_cache_key on requests to the same model that share a prompt prefix. This routing hint also applies to ASAP requests, with or without streaming. Available capacity and current load can override the hint, so it does not guarantee a cache hit.

Supercache usage

When Supercache accounting data is available, completed Responses include metadata.supercached_input_tokens and metadata.supercache_write_input_tokens as decimal strings. Both fields are included when their value is zero. usage.input_tokens_details.cached_tokens includes tokens read from Supercache. Both metadata keys are reserved and cannot be set in a request.

Response status and output validation

Treat completed, incomplete, failed, and cancelled as terminal statuses when polling a background response.
  • completed means the response finished normally.
  • incomplete means generation stopped early. A reason of "max_output_tokens" is Sail’s normalized truncation/cap reason: it is reported whenever generation stops because of a length-based limit, such as the request’s max_output_tokens, the model’s context window, or another provider length signal. It does not by itself prove the request’s output limit alone was reached. A reason of "content_filter" means generation stopped because of a content filter. When present, usage is preserved. A response limited by max_output_tokens can include partial output. Tokens reported in usage are billed normally. An incomplete response is not an API error.
  • failed includes generations that Sail rejects because the result does not satisfy the request after its retry attempts are exhausted. A request with tool_choice: "required" can fail with error.code: "server_error" and the message "The model did not return a tool call required by the request." A generation with no visible text, refusal, or tool call can fail with error.code: "server_error" and the message "The model did not return visible text, a refusal, or a tool call." instead of returning an empty completed response.
  • cancelled means no further output will be produced.
Sail cannot currently guarantee tool_choice: "required" for openai/gpt-oss-* models. Choose another model when every successful response must contain a tool call.

Not yet supported


Chat Completions API

OpenAI SDK compatibleAPI reference

Supported features

Not yet supported

Response notes

  • Responses always contain exactly one choice (n=1).
  • finish_reason reflects the provider result when available, including "stop", "tool_calls", "length", and "content_filter".
  • system_fingerprint and service_tier are not included in responses.
  • logprobs is always null.

Messages API

Anthropic Messages formatAnthropic SDK compatibleAPI reference
The Messages API is Anthropic-compatible for agentic use: system prompts, tool calling, and streaming (SSE) are supported. Prompt caching is automatic, so you don’t need cache_control breakpoints.

Supported features

Not yet supported

Response notes

  • stop_reason reflects the outcome. Sail returns "end_turn" normally, "tool_use" for tool calls, "max_tokens" for token limits, and "refusal" when the provider reports a refusal. Sail returns "model_context_window_exceeded" when the provider reports that the model’s context window was exceeded.
  • Responses contain text content blocks, plus tool_use blocks when the model calls a tool.
  • usage follows Anthropic’s accounting. input_tokens counts only the part of the prompt that was not read from cache, and cache_read_input_tokens counts the part that was. The full prompt is their sum. cache_creation_input_tokens is always 0: Sail has no separate charge for writing the cache, so every uncached prompt token is in input_tokens.
  • Thinking output does not include an Anthropic cryptographic signature. thinking.budget_tokens is approximated as medium reasoning effort unless output_config.effort provides an explicit effort.
  • thinking.type: "disabled" uses the selected model’s none reasoning control. Sail returns 400 when the selected model does not support disabling reasoning. The exact effect follows that model’s none contract, so disabled does not bypass a model-specific lowest-effort approximation. output_config.effort overrides the effort level. A non-disabling override preserves thinking output requested by thinking.type: "enabled" or "adaptive". An effort whose model mapping disables reasoning can remove it.
  • When reasoning is present, Messages content begins with a thinking block before the text block. Select content by block type rather than assuming content[0] is text. The Chat Completions response exposes the same reasoning through reasoning_content.
  • Token counting applies the same replay-block policy as message creation: redacted thinking is omitted and tool-result images count as the replacement text marker.

Compatibility notes

  • Sail accepts both the Anthropic x-api-key header and Authorization: Bearer <key>. If both are present, Authorization takes precedence.
  • The Anthropic Python and TypeScript SDKs type metadata with only user_id. Sail also accepts completion_window. Keep the additional cast or type assertion scoped to the metadata value:
  • The anthropic-version header is not required or checked.
  • Errors use the Anthropic envelope {"type":"error","error":{"type":"...","message":"..."},"request_id":"..."} and Anthropic error types such as invalid_request_error, rate_limit_error, and overloaded_error.

Batch API

The Batch API runs large numbers of Responses API requests asynchronously. Every item targets /v1/responses. Batching /v1/chat/completions or /v1/messages is not currently supported. You can submit up to 100,000 requests in one POST /v1/batches call, then poll GET /v1/batches/{id} for progress and inline results when the batch finishes. For batches above the inline size limit, fetch results by custom_id. See Sending requests at scale for the end-to-end workflow and the Batch API reference for the request and response schemas.

Cross-API behavior

These behaviors apply across the inference API surfaces:
  • Streaming: The Chat Completions API supports stream: true, returning Server-Sent Events (chat.completion.chunk). Set stream_options.include_usage for a final usage chunk. The Messages API supports stream: true, returning Anthropic SSE events as output becomes available. The Responses API supports foreground stream: true, returning OpenAI Responses SSE events. Accepted background requests return 202 immediately and cannot be streamed. Use polling or webhooks for long-running background work. A queued foreground request can stream once generation starts. Some request types deliver content only after completion. Streaming does not yet continue across a failed execution attempt.
  • Completion windows: Express latency tolerance in exchange for lower token costs. See Flex request requirements and Pricing.
  • Context windows: For each text request, Sail reserves at least 512 tokens, or 0.5% on larger windows, from the model’s published context window. The input token count plus the requested maximum output must fit in the remaining budget. This allowance covers model-specific request formatting that can add tokens when the request runs. On the non-batch Responses endpoint (POST /v1/responses), requests that provide raw_prompt_tokens skip text formatting and use the exact raw-token count without this reserve, requiring the raw token count plus max_output_tokens to fit within the model’s context window. The Batch API does not yet apply this raw_prompt_tokens admission arithmetic.
  • Webhooks: Set metadata.completion_webhook to receive a POST when processing finishes. See Webhooks.
  • Response storage: store: false is accepted for OpenAI compatibility, but does not change Sail’s normal temporary request/response storage for processing, retries, polling, and idempotency. Customer Data remains governed by Sail’s DPA retention and deletion terms.

PDF input

Sail reads the text of each page of a PDF and sends that text to the model in place of the file. PDF input therefore works on every model, including models that accept only text. The model does not see page images, so charts, figures, and scanned pages are not visible to it. Each page’s text is labeled with its page number. When the request includes a filename (Responses and Chat Completions) or title (Messages), the PDF is labeled with it as well. The extracted text counts toward input tokens and the model’s context window, and token counting endpoints count it the same way. These requests are rejected with HTTP 400:
  • A password-protected PDF.
  • A PDF with no text layer, such as a scanned document. Send its pages as images to a multimodal model instead.
  • A PDF with more than 1,000 pages.
  • More than 50 MB of PDF files in one request.
  • More than 16 MB of extracted text in one request.
  • A PDF referenced by file ID or URL. Send the file inline as base64.
  • Any file input in a Batch API request.