Batch API
Submit inference requests in a batch, then pollGET /v1/batches/{batch_id} until status is completed. The final response includes a results array, with each result matched to your input by custom_id.
Use an idempotency key when submitting so retries don’t create duplicate batches. To list your batches, use GET /v1/batches.
Batch requests always use the
flex completion window.
For low-latency workloads, use the Responses
API instead.Python example
First, install therequests library:
Reading the batch response
statusisin_progressuntil every request finishes, thencompleted, even if some requests failed or were cancelled.request_countsshows how many requests were submitted, completed, or failed.request_statuslists each request’s status bycustom_id. This list may be incomplete while the batch is running. Stop polling when the batch’s top-levelstatusiscompleted.- Each result has either an
erroror aresponse.bodycontaining a Responses API object. Check the body’s status for incomplete output, and match results bycustom_id.
Large batches and individual results
- Submit up to 100,000 requests per batch, with a maximum request body of 256 MiB.
- Inline results are limited to approximately 1 GiB per batch. If the results are too large, the response tells you to retrieve them using the individual batch results endpoint (see the reference code below).
- Inline results are one JSON response. If the connection ends before it finishes, retry the GET.
Responses API with background mode
You can also submit requests individually using the Responses API withbackground=True and setting the completion window to balanced or flex. For large-volume workloads (1,000+ requests), we recommend:
- Use
AsyncOpenAIwithDefaultAioHttpClient(). The OpenAI SDK’s built-in aiohttp client is more efficient than the defaulthttpxbackend for high-concurrency workloads. - Limit concurrency with an
asyncio.Semaphore. The semaphore caps how many connections your client has open at once (such as 200), so the machine sending the requests does not run out of connections. - Submit all requests concurrently with
background=True, then poll. Send all submissions in parallel (limited by the semaphore), collect the response IDs, and poll for completions in a separate loop. - Use one idempotency key per submission. Reuse it if you need to retry that submission.