> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sailresearch.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Handle request limits and temporary capacity shortages

Sail limits how quickly your organization can send inference requests and how
many can run at once. API keys in the same organization share these limits.

Available capacity can vary during periods of high demand. Pro and Enterprise
customers receive increased access to model capacity and higher rate limits,
with Enterprise offering the highest capacity.

## 429: Too many requests

Your organization has reached a request or concurrency limit. Send fewer
requests or run fewer at once. Wait for the delay in `Retry-After` before
retrying.

## 503: Service unavailable

Sail cannot serve the request right now. Unavailable model capacity is one
possible cause. Some `503` errors can occur after Sail has accepted the
request. Follow [Retrying requests](/retries) before submitting again.

When a response includes `Retry-After`, its value is the minimum delay in
seconds before retrying.

## Request higher limits

[Upgrade to Pro](/pricing#plans) for higher limits and prioritized access to
models. For the highest capacity, or if your workload has specific throughput
needs, [contact support](https://www.sailresearch.com/support) to ask about
Enterprise.
