Skip to main content
For large workloads, submit a batch or send concurrent Responses API requests in background mode.

Batch API

Submit inference requests in a batch, then poll GET /v1/batches/{batch_id} until status is completed. The final response includes a results array, with each result matched to your input by custom_id. Use an idempotency key when submitting so retries don’t create duplicate batches. To list your batches, use GET /v1/batches.
Batch requests always use the flex completion window. For low-latency workloads, use the Responses API instead.

Python example

First, install the requests library:

Reading the batch response

  • status is in_progress until every request finishes, then completed, even if some requests failed or were cancelled.
  • request_counts shows how many requests were submitted, completed, or failed.
  • request_status lists each request’s status by custom_id. This list may be incomplete while the batch is running. Stop polling when the batch’s top-level status is completed.
  • Each result has either an error or a response.body containing a Responses API object. Check the body’s status for incomplete output, and match results by custom_id.
See the endpoint reference for the full response shape.

Large batches and individual results

  • Submit up to 100,000 requests per batch, with a maximum request body of 256 MiB.
  • Inline results are limited to approximately 1 GiB per batch. If the results are too large, the response tells you to retrieve them using the individual batch results endpoint (see the reference code below).
  • Inline results are one JSON response. If the connection ends before it finishes, retry the GET.
You can also fetch individual results as soon as they finish:

Responses API with background mode

You can also submit requests individually using the Responses API with background=True and setting the completion window to balanced or flex. For large-volume workloads (1,000+ requests), we recommend:
  1. Use AsyncOpenAI with DefaultAioHttpClient(). The OpenAI SDK’s built-in aiohttp client is more efficient than the default httpx backend for high-concurrency workloads.
  2. Limit concurrency with an asyncio.Semaphore. The semaphore caps how many connections your client has open at once (such as 200), so the machine sending the requests does not run out of connections.
  3. Submit all requests concurrently with background=True, then poll. Send all submissions in parallel (limited by the semaphore), collect the response IDs, and poll for completions in a separate loop.
  4. Use one idempotency key per submission. Reuse it if you need to retry that submission.

Python example

First, install tqdm and the OpenAI SDK with the aiohttp extra:
The semaphore limit of 200 is a good starting point. Lower it if you run into connection errors or timeouts. Raise it if your machine can hold more open connections and you want faster submission throughput.