Skip to main content
Completion windows let you express how long your requests can wait, giving Sail room to increase cost efficiency. Sail serves low-latency inference by default for core models, at lower cost than traditional inference providers. You can opt in to the balanced or flex completion windows to cut token costs drastically, for work that is latency tolerant.

Completion windows at a glance

Check Pricing for current per-model availability.

Completion window details

asap

asap is the default, low-latency path for Sail’s core models.
asap is not compatible with Responses API requests with background=True or Batch API requests.

balanced

balanced gives Sail more time to place work on efficient capacity, greatly lowering token costs for long-horizon background agents and pipelines.

flex

flex gives Sail the widest scheduling window and offers the lowest available token prices. Common use cases include batch jobs, evals, and offline processing.
flex is not compatible with Chat Completions, Messages, or foreground (background=False) Responses API requests.

How to set completion windows

Via request body metadata

Set metadata.completion_window to asap, balanced, or flex:
To use flex, you must use the Responses API with background=True or submit requests through the Batch API.
flex is not compatible with Chat Completions, Messages, or foreground (background=False) Responses API requests.

Via request header

Alternatively, if your client cannot add metadata to a request body, you can set the completion window using the X-Sail-Completion-Window request header.

Default behavior

For inference requests for core models, the default behavior (when no completion_window is specified) is low-latency inference (asap). The exceptions to this default behavior are:
  • Batch inference requests, which default to flex and do not support other completion windows.
  • LoRA requests, which default to balanced for core models.
  • Requests that use custom tools, which default to balanced for core models because asap does not support them.
  • Responses API requests with background=True use balanced or flex by default, since they do not support asap.
For flex-only models, the default (and only) completion window is flex.

Troubleshooting

asap requests

  • asap does not support custom tools (tools with type="custom") on the Responses or Chat Completions APIs. Use balanced, or omit the completion window to default to balanced for core models.
  • asap does not support Responses requests with background=True or the Batch API. Use balanced or flex for Responses requests with background=True.
  • Setting metadata.completion_window="asap" with custom tools or background=True returns HTTP 400 with error code unsupported_asap_request.

balanced requests

  • balanced requests wait in a queue until Sail can run them. You can send them synchronously or as Responses requests with background=True.
  • A synchronous streaming balanced request opens its stream right away and keeps the connection alive while it waits. Output starts when the model begins generating.

flex requests

  • Chat Completions and Messages reject flex with HTTP 400, with or without streaming. Use asap or balanced on those endpoints.
  • Responses requests using flex reject an omitted background value, background=None, or background=False with HTTP 400. Set background=True explicitly.
  • Sail queues accepted flex requests until capacity is available. During busy periods, they may wait a long time.

Request settings

  • If you set both X-Sail-Completion-Window and metadata.completion_window, the value in the request body takes precedence.
  • The model must support the window you choose. Check Pricing for availability.

Batch and Background

  • Responses requests with background=True cannot use stream=True. Retrieve results by polling or using webhooks.
  • A polling timeout does not cancel a Responses request with background=True. Keep polling the same response ID. See Retrying requests if you lose the connection before receiving an ID.
  • The Batch API does not use X-Sail-Completion-Window.
  • Batch always uses flex and runs in the background. Setting a subrequest’s completion window to asap or balanced returns HTTP 400.
  • Batch rejects models that do not support flex.

Capacity errors

  • For asap requests, Sail waits briefly for capacity. If none becomes available, it returns 503 with code model_capacity_unavailable without queueing the request.
  • If the service is overloaded, Sail can return 503 with code service_overloaded before accepting a request in any completion window.
  • If the response includes Retry-After, wait at least that many seconds. See Retrying requests before resubmitting after a timeout or another server error, since the request may already have started.