Responses API
Supported features
Prompt cache routing
For Responses and Chat Completions, reuse the sameprompt_cache_key on requests
to the same model that share a prompt prefix. This routing hint also applies to
ASAP requests, with or without streaming. Available capacity and current load can
override the hint, so it does not guarantee a cache hit.
Supercache usage
When Supercache accounting data is available, completed Responses includemetadata.supercached_input_tokens and
metadata.supercache_write_input_tokens as decimal strings. Both fields are
included when their value is zero. usage.input_tokens_details.cached_tokens
includes tokens read from Supercache. Both metadata keys are reserved and
cannot be set in a request.
Response status and output validation
Treatcompleted, incomplete, failed, and cancelled as terminal statuses
when polling a background response.
completedmeans the response finished normally.incompletemeans generation stopped early. A reason of"max_output_tokens"is Sail’s normalized truncation/cap reason: it is reported whenever generation stops because of a length-based limit, such as the request’smax_output_tokens, the model’s context window, or another provider length signal. It does not by itself prove the request’s output limit alone was reached. A reason of"content_filter"means generation stopped because of a content filter. When present,usageis preserved. A response limited bymax_output_tokenscan include partialoutput. Tokens reported inusageare billed normally. An incomplete response is not an API error.failedincludes generations that Sail rejects because the result does not satisfy the request after its retry attempts are exhausted. A request withtool_choice: "required"can fail witherror.code: "server_error"and the message"The model did not return a tool call required by the request."A generation with no visible text, refusal, or tool call can fail witherror.code: "server_error"and the message"The model did not return visible text, a refusal, or a tool call."instead of returning an empty completed response.cancelledmeans no further output will be produced.
Not yet supported
Chat Completions API
OpenAI SDK compatibleAPI reference
Supported features
Not yet supported
Response notes
- Responses always contain exactly one choice (
n=1). finish_reasonreflects the provider result when available, including"stop","tool_calls","length", and"content_filter".system_fingerprintandservice_tierare not included in responses.logprobsis alwaysnull.
Messages API
The Messages API is Anthropic-compatible for agentic use: system prompts, tool
calling, and streaming (SSE) are supported. Prompt caching is automatic, so
you don’t need
cache_control breakpoints.Supported features
Not yet supported
Response notes
stop_reasonreflects the outcome. Sail returns"end_turn"normally,"tool_use"for tool calls,"max_tokens"for token limits, and"refusal"when the provider reports a refusal. Sail returns"model_context_window_exceeded"when the provider reports that the model’s context window was exceeded.- Responses contain
textcontent blocks, plustool_useblocks when the model calls a tool. usagefollows Anthropic’s accounting.input_tokenscounts only the part of the prompt that was not read from cache, andcache_read_input_tokenscounts the part that was. The full prompt is their sum.cache_creation_input_tokensis always0: Sail has no separate charge for writing the cache, so every uncached prompt token is ininput_tokens.- Thinking output does not include an Anthropic cryptographic signature.
thinking.budget_tokensis approximated as medium reasoning effort unlessoutput_config.effortprovides an explicit effort. thinking.type: "disabled"uses the selected model’snonereasoning control. Sail returns400when the selected model does not support disabling reasoning. The exact effect follows that model’snonecontract, sodisableddoes not bypass a model-specific lowest-effort approximation.output_config.effortoverrides the effort level. A non-disabling override preserves thinking output requested bythinking.type: "enabled"or"adaptive". An effort whose model mapping disables reasoning can remove it.- When reasoning is present, Messages content begins with a
thinkingblock before the text block. Select content by block type rather than assumingcontent[0]is text. The Chat Completions response exposes the same reasoning throughreasoning_content. - Token counting applies the same replay-block policy as message creation: redacted thinking is omitted and tool-result images count as the replacement text marker.
Compatibility notes
- Sail accepts both the Anthropic
x-api-keyheader andAuthorization: Bearer <key>. If both are present,Authorizationtakes precedence.
- The Anthropic Python and TypeScript SDKs type
metadatawith onlyuser_id. Sail also acceptscompletion_window. Keep the additional cast or type assertion scoped to themetadatavalue:
- The
anthropic-versionheader is not required or checked. - Errors use the Anthropic envelope
{"type":"error","error":{"type":"...","message":"..."},"request_id":"..."}and Anthropic error types such asinvalid_request_error,rate_limit_error, andoverloaded_error.
Batch API
The Batch API runs large numbers of Responses API requests asynchronously. Every item targets/v1/responses. Batching /v1/chat/completions or /v1/messages is not currently supported. You can submit up to 100,000 requests in one POST /v1/batches call, then poll GET /v1/batches/{id} for progress and inline results when the batch finishes. For batches above the inline size limit, fetch results by custom_id.
See Sending requests at scale for the end-to-end workflow and the Batch API reference for the request and response schemas.
Cross-API behavior
These behaviors apply across the inference API surfaces:- Streaming: The Chat Completions API supports
stream: true, returning Server-Sent Events (chat.completion.chunk). Setstream_options.include_usagefor a final usage chunk. The Messages API supportsstream: true, returning Anthropic SSE events as output becomes available. The Responses API supports foregroundstream: true, returning OpenAI Responses SSE events. Accepted background requests return202immediately and cannot be streamed. Use polling or webhooks for long-running background work. A queued foreground request can stream once generation starts. Some request types deliver content only after completion. Streaming does not yet continue across a failed execution attempt. - Completion windows: Express latency tolerance in exchange for lower token costs. See Flex request requirements and Pricing.
- Context windows: For each text request, Sail reserves at least 512 tokens,
or 0.5% on larger windows, from the model’s published context window. The
input token count plus the requested maximum output must fit in the remaining
budget. This allowance covers model-specific request formatting that can add
tokens when the request runs. On the non-batch Responses endpoint
(
POST /v1/responses), requests that provideraw_prompt_tokensskip text formatting and use the exact raw-token count without this reserve, requiring the raw token count plusmax_output_tokensto fit within the model’s context window. The Batch API does not yet apply thisraw_prompt_tokensadmission arithmetic. - Webhooks: Set
metadata.completion_webhookto receive a POST when processing finishes. See Webhooks. - Response storage:
store: falseis accepted for OpenAI compatibility, but does not change Sail’s normal temporary request/response storage for processing, retries, polling, and idempotency. Customer Data remains governed by Sail’s DPA retention and deletion terms.
PDF input
Sail reads the text of each page of a PDF and sends that text to the model in place of the file. PDF input therefore works on every model, including models that accept only text. The model does not see page images, so charts, figures, and scanned pages are not visible to it. Each page’s text is labeled with its page number. When the request includes afilename (Responses and Chat Completions) or title (Messages), the PDF is
labeled with it as well. The extracted text counts toward input tokens and the
model’s context window, and token counting endpoints count it the same way.
These requests are rejected with HTTP 400:
- A password-protected PDF.
- A PDF with no text layer, such as a scanned document. Send its pages as images to a multimodal model instead.
- A PDF with more than 1,000 pages.
- More than 50 MB of PDF files in one request.
- More than 16 MB of extracted text in one request.
- A PDF referenced by file ID or URL. Send the file inline as base64.
- Any file input in a Batch API request.