Common parameters
Most routes accept a rollingrange and, for the operational routes, an
environment filter.
Relative billing ranges count backward from the current time. Billing data for
1h, 6h, and 24h is grouped by hour. For 7d, 30d, and period, partial
days at the beginning and end use hourly groups, while complete days use daily
groups. Billing totals may extend to the hour boundaries just before and after
the selected range. Activity and latency data use the exact times selected.
range never returns a 400 for an unrecognized value. It silently falls back
to the endpoint’s default range. Invalid environment, date, start,
end, bucket_size, status, kind, or after values do return 400, as
does an unsupported sla on /tasks and /latency/timeseries. On GET /v2/usage, an unrecognized sla or model is not rejected. It is applied as
a filter that doesn’t match anything.Your plan limits how far back you can read.- Free organizations can read up to 7 days back and see the last hour of
request activity. Explicit
startvalues andrange=daydates are floored at 7 days ago. - Pro organizations can read up to 30 days back and see the last 24 hours of
request activity. Explicit
startvalues andrange=daydates are floored at 30 days ago. - Enterprise organizations can select rolling ranges up to 30 days or the
current billing
period, see the last 24 hours of request activity, and read any explicitstartvalue orrange=daydate, subject to retention.
Monetary fields (
balance, period_spend, burn_rate, avg_cost_per_day,
product_spend, sailbox_spend, and breakdown total) are fractional USD
cents expressed as floats. A value of 73795.09 is roughly $737.95. Token
fields on the public endpoints are raw inference token counts.Product accounting
Billing responses follow these guarantees:period_spendand each breakdown bucket’stotalinclude all positive Inference and Sailbox charges.product_spendalways contains exactly two product families,inferenceandsailboxes. It never emits anotherfamily, and the two values add up to the corresponding combined total.sailbox_spendreports the currently itemized Sailbox charges. Sailbox products without a line item still count towardproduct_spend.sailboxes, so the line-item fields may sum to less than it.- Token, model, completion-window, request, and latency metrics are inference-only. Sailbox quantities and identifiers are not represented as tokens, models, or completion windows.
- Only base input, output, and cached-input token products add token counts.
Surcharges add inference spend without duplicating token quantities. Cached
input is included in
input, sototal = input + outputand cached tokens must not be added tototalagain. - Model and completion-window breakdowns include only explicitly attributed
inference spend. Missing attribution stays unassigned instead of creating an
othermodel or completion window.
Billing & cost
GET /v2/usage/summary
Combined Inference and Sailbox spend, credit balance, burn rate, days remaining, inference token totals, inference SLA spend mix, and a prior-period comparison.
Both windows contain UTC
start and end timestamps. The start is inclusive
and the end is exclusive. These describe the query window, not how recently
usage data has updated.
GET /v2/usage/breakdown
Combined spend and inference token breakdowns per time bucket, plus a ranking of inference models by cost. Userange=day with date=YYYY-MM-DD to get
hourly buckets for one day.
With
range=day, a date older than your plan allows returns an
empty breakdown with plan_limited: true rather than an error. See
Common parameters.
Inference spend without a completion-window or tool attribution is the
remainder below. It is plain inference spend, not an
other product or
completion window.
balanced window under the key standard.
GET /v2/usage/api-keys
Per-API-key usage, broken down by(api_key_id, model, sla), with a
time series. display_name and display_prefix are populated for keys that
still exist. Deleted keys return null for both, and only the stable
api_key_id remains.
This endpoint reports inference usage only. Sailbox charges accrue over a
Sailbox’s lifetime and do not map reliably to one API key, so Sailbox usage is
excluded from the per-key rows. Use
/v2/usage/summary or
/v2/usage/breakdown for product-aware spend.GET /v2/usage/tokens
Inference token counters (input/output/cached) over the range, as raw counts. Only base token products contribute quantities. Sailbox usage and inference surcharges do not add tokens. Cached input is already included ininput.
GET /v2/usage/tokens/timeseries
Raw inference token counts per time bucket, broken down by input/output/cached. The same base-product and cached-input rules as/tokens apply.
Operational activity & latency
These routes report request counts, task activity, and latency. They all accept theenvironment filter.
GET /v2/usage/activity
Consolidated completed-request counts and average latency over rolling ranges.GET /v2/usage/activity/timeseries
Per-model completed-request counts over time.GET /v2/usage/recent
The most recent requests, including active and finished requests when available.sla is null when the request has no resolved completion window. The listing
covers only your plan’s request-log range: the last hour on Free, the last 24
hours on Pro and Enterprise. Older requests are not listed even when they are
still retained.
GET /v2/usage/tasks
Paginated task activity with per-task token breakdowns, spanning both active and finished requests.
Only
1h and 24h are valid for range. Any other value resolves to 24h,
or to 1h for Free organizations (reported back as effective_range).
GET /v2/usage/latency/turn
Turn latency distribution for a single request-to-response time, with p50/p95/p99, aggregate and per-SLA.GET /v2/usage/latency/trajectory
Trajectory latency distribution for a multi-turn conversation, with the same response format as/latency/turn.
Trajectory responses also include avg_turns_per_trajectory, and may set
approx: true when percentiles are sampled.
GET /v2/usage/latency/timeseries
Latency percentiles over time for turns or trajectories.model and sla filters are only supported for kind=turn. Passing either
with kind=trajectory returns a 400.approx and avg_turns_per_trajectory only when relevant
(both are omitted otherwise).
GET /v2/usage
Completed-request latency and counts in fixed time buckets, with one row per environment per bucket. Unlike the rolling-range routes above, this endpoint takes explicitstart, end, and bucket_size parameters. Use it to chart a
specific time range at a fixed resolution.
A range that would produce more than 10,000 buckets for the chosen
bucket_size returns a 400. Widen bucket_size or shorten the range.Errors
Errors use the same envelope as the inference API. For usage-API errors,code
is set to the same string as type, and param is always null:
code (such as invalid_api_key) or null (for a
missing Authorization header).
Endpoints return an empty/zeroed payload with HTTP 200 (not an error) when the
org has no billing account or no usage in the requested range.