balanced or flex completion windows to cut token costs drastically, for work that is latency tolerant.
Completion windows at a glance
Check Pricing for current per-model availability.
Completion window details
asap
asap is the default, low-latency path for Sail’s core models.
balanced
balanced gives Sail more time to place work on efficient capacity, greatly
lowering token costs for long-horizon background agents and pipelines.
flex
flex gives Sail the widest scheduling window and offers the lowest available
token prices. Common use cases include batch jobs, evals, and offline processing.
How to set completion windows
Via request body metadata
Set metadata.completion_window to asap, balanced, or flex:
flex, you must use the Responses API
with background=True or submit requests through the Batch API.
Via request header
Alternatively, if your client cannot add metadata to a request body, you can set the completion window using theX-Sail-Completion-Window request header.
Default behavior
For inference requests for core models, the default behavior (when nocompletion_window is specified) is low-latency inference (asap).
The exceptions to this default behavior are:
- Batch inference requests, which default to
flexand do not support other completion windows. - LoRA requests, which default to
balancedfor core models. - Requests that use custom tools, which default to
balancedfor core models becauseasapdoes not support them. - Responses API requests with
background=Trueusebalancedorflexby default, since they do not supportasap.
flex.
Troubleshooting
asap requests
asapdoes not support custom tools (toolswithtype="custom") on the Responses or Chat Completions APIs. Usebalanced, or omit the completion window to default tobalancedfor core models.asapdoes not support Responses requests withbackground=Trueor the Batch API. Usebalancedorflexfor Responses requests withbackground=True.- Setting
metadata.completion_window="asap"with custom tools orbackground=Truereturns HTTP400with error codeunsupported_asap_request.
balanced requests
balancedrequests wait in a queue until Sail can run them. You can send them synchronously or as Responses requests withbackground=True.- A synchronous streaming
balancedrequest opens its stream right away and keeps the connection alive while it waits. Output starts when the model begins generating.
flex requests
- Chat Completions and Messages reject
flexwith HTTP400, with or without streaming. Useasaporbalancedon those endpoints. - Responses requests using
flexreject an omittedbackgroundvalue,background=None, orbackground=Falsewith HTTP400. Setbackground=Trueexplicitly. - Sail queues accepted
flexrequests until capacity is available. During busy periods, they may wait a long time.
Request settings
- If you set both
X-Sail-Completion-Windowandmetadata.completion_window, the value in the request body takes precedence. - The model must support the window you choose. Check Pricing for availability.
Batch and Background
- Responses requests with
background=Truecannot usestream=True. Retrieve results by polling or using webhooks. - A polling timeout does not cancel a Responses request with
background=True. Keep polling the same response ID. See Retrying requests if you lose the connection before receiving an ID. - The Batch API does not use
X-Sail-Completion-Window. - Batch always uses
flexand runs in the background. Setting a subrequest’s completion window toasaporbalancedreturns HTTP400. - Batch rejects models that do not support
flex.
Capacity errors
-
For
asaprequests, Sail waits briefly for capacity. If none becomes available, it returns503with codemodel_capacity_unavailablewithout queueing the request. -
If the service is overloaded, Sail can return
503with codeservice_overloadedbefore accepting a request in any completion window. -
If the response includes
Retry-After, wait at least that many seconds. See Retrying requests before resubmitting after a timeout or another server error, since the request may already have started.