> ## Documentation Index
> Fetch the complete documentation index at: https://docs.phoneuse.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Three budgets, one refusal shape, and the headers that tell you where you stand.

Every request to the hosted HTTP surfaces passes a limiter before it reaches a
route. The limiter sits ahead of routing and ahead of authentication, so being
refused here tells you nothing about whether your credential was any good. One
thing is answered before the limiter and never counted: a browser's CORS
preflight. See [If you are calling from a browser](#if-you-are-calling-from-a-browser).

The numbers on this page are **defaults**. Every response that a budget applied
to carries that budget in its headers, and those headers are authoritative: read
them rather than hard coding what you find here.

## The three budgets

| | Run | Credentialed | Bare |
| - | - | - | - |
| Applies to | A request on `/vrx/{runToken}/v1/...` naming a live run | Any other request with an `Authorization` header | Everything else |
| Default budget | 600 requests / minute | 300 requests / minute | 60 requests / minute |
| Counted per | The run | The credential | The client IP |

Outside the device surface, which budget applies is decided by one thing:
whether **this request** carries an `Authorization` header. It is a property of
the request, not of the endpoint you aimed it at, so a client that adds a header
from shared configuration changes its own budget.

On the device surface the run wins, and the header is ignored there for pacing
exactly as it is ignored for authorization.

Two things follow from "counted per credential" that are easy to get wrong:

* **The budget is per credential, not per org.** Two API keys in the same org
  hold two independent budgets, and one key exhausting its own cannot slow the
  other down. Respelling the header does not buy you a second budget: the
  credential is identified the same way the authenticator identifies it.
* **A rotated key starts fresh.** A new token is a new budget. The limiter keeps
  only a hash, never the credential itself.

There is also a ceiling on how many distinct credentials one client IP may hold
separate budgets for at once. Past it, a credential that address has not been
using is counted against its bare budget instead, and a slot is released once
its credential goes quiet. Reaching the ceiling takes more than a handful of
credentials in simultaneous use from one address, so if that is your shape, say
so when you ask for access. It exists so that presenting a credential cannot
itself become a way around the bare budget.

### If you are calling from a browser

A browser sends a preflight (`OPTIONS`) before any request that carries an
`Authorization` header. On a deployment that has granted your origin, the
preflight is answered **before** the limiter and is not counted against any
budget, yours or the bare one. It carries no `RateLimit-*` headers, because
none applied. On a deployment with no origins granted there is no preflight
handling at all, and an `OPTIONS` is counted like any other bare request.

That is deliberate. A preflight never carries a credential, so counting it could
only put it in the bare budget, shared with every other caller behind your
egress address. And a refused preflight would never reach your code as a `429`:
browsers treat any non-2xx preflight as a network error, with no status and no
headers to read. So the API never refuses a preflight for rate, and a CORS
error in your browser console is a CORS problem, not a rate limit.

The request that follows the preflight is counted against your credential in the
ordinary way. If it is refused, you get a real `429`: the response carries
`Access-Control-Allow-Origin`, and `Retry-After` and the `RateLimit-*` headers
are exposed, so your code can read them and back off.

One caveat that applies to browsers more than to anything else: a session token
that rotates presents a **new credential** on every rotation, and the ceiling on
distinct credentials per client IP described above is reached by rotation
alone. Past it, requests are counted against the bare budget until an older
credential goes quiet, so a browser can find itself on the smaller budget
without having changed anything it does.

For bare requests the client IP is the one our edge recorded for the
connection, falling back to the first hop in `X-Forwarded-For`.

### The device surface

The device surface, `/vrx/{runToken}/v1/...`, carries its credential in the URL
path and ignores any `Authorization` header, by design. It gets its own budget,
counted **per run**.

That means two runs behind one egress address do not share an allowance, and
one run spending its own cannot slow another down. Adding workers behind the
same address adds capacity rather than dividing it.

The budget applies to a run the control plane currently holds live. A request
naming a run token that has expired, ended, or never existed is counted against
the **bare** budget for your client IP, alongside every other uncredentialed
caller behind that address. That is the same answer such a request has always
had, so a worker still calling after its run is over pays the bare rate.

The switch happens when the run ends, not when you next call, so there is no
window in which a finished run is still spending at the run rate.

<Note>
  A run's budget bounds how fast it may call, not how much it may do in total.
  The total is the run's step budget, which is fixed when the run starts and
  enforced before any frame is issued. Raising one does not raise the other.
</Note>

Two consequences worth planning around:

* **The budget arrives with the run.** It is in force from your first request,
  so there is nothing to register and nothing to wait for.
* **It does not survive the run.** The moment the run reaches a terminal state
  its token stops working and its budget is withdrawn, so the next call gets a
  `404` counted against the bare budget.

## Being refused

Over budget is `429`, with the control plane's HTTP error envelope:

```json theme={null}
{ "error": { "code": "rate_limited", "message": "Too many requests. Slow down and retry." } }
```

This is the HTTP envelope, not the MCP tool error envelope. The limiter runs
before the MCP endpoint is reached, so a throttled tool call comes back as a
transport level `429`, **not** as a result with `isError` and a `code` from the
[error reference](/mcp/errors). A client that only branches on tool errors will
see this as a transport failure, which is what it is.

`Retry-After` on that response is an integer number of seconds, computed from
how far under one request the caller actually is. It is not a fixed backoff
constant, so honour it rather than replacing it with your own.

## The headers

Every response that a budget applied to carries these:

| Header | Value |
| - | - |
| `RateLimit-Limit` | The budget in effect for this caller class, in requests per minute |
| `RateLimit-Remaining` | Whole requests still available when this request was admitted |
| `RateLimit-Reset` | Seconds from that moment until the allowance is back to full |
| `RateLimit-Policy` | The budget and its window, as `<limit>;w=60` |

Both counts describe the allowance at the moment the request was admitted, not
at the moment you read them. On a long running call the real figure has already
crept up, so they understate rather than overstate what you have left.

`RateLimit-Reset` and `Retry-After` answer different questions and will differ.
Reset is seconds until **full**. `Retry-After` on a `429` is seconds until
**one** request is available. Pacing against Reset when you only need one call
will make you wait longer than necessary.

The allowance refills continuously rather than resetting on a boundary, so a
caller that has been idle can spend up to the full budget at once and then
drains toward the steady rate.

**Absence is information.** No `RateLimit-*` headers on a response means no
budget applied to it, either because the path is exempt or because that caller
class is switched off in this deployment. It never means "unknown".

## What is exempt

`GET` and `HEAD` on `/healthz` are never limited, in any caller class. A `429`
on a health probe would read as an outage and mask the real ones. Only those two
methods are exempt, and only that exact path.

On a deployment that has granted at least one origin, a CORS preflight on a
path a browser may reach (`/v1/*`, `/healthz` and the discovery
documents) is answered before the limiter and never counted. Only a real
preflight qualifies: an `OPTIONS` request that carries an `Origin` header. An
`OPTIONS` without one is an ordinary request and is counted like any other.

## What the budgets are scoped to

Two boundaries worth knowing before you plan around these numbers.

* **Per serving process.** The allowance lives in the process that answers you.
  Read the numbers as a floor you can rely on rather than a global ceiling on
  the deployment, and pace against the headers you actually receive.
* **Hosted only.** Local mode has no limiter at all. It binds to loopback and
  serves one operator, so there is nothing to protect and no headers to read. Do
  not use local behaviour to predict hosted behaviour here.

## Living within them

* Read `RateLimit-Remaining` and slow down before you are refused. Being refused
  costs you a round trip and tells you nothing you could not have read from the
  previous response.
* On a `429`, wait `Retry-After` seconds and retry the same request. The refusal
  happens before any work, so nothing was half applied and the retry is safe.
* Do not retry a `429` immediately or in a tight loop. The allowance refills on
  a clock, and retrying faster cannot make it refill faster.
* If your steady state does not fit the default budget, say so when you ask for
  access. The budget is a deployment setting, not a property of the protocol.

### If you are driving a device

One rule matters more on this surface than anywhere else, because getting it
wrong is expensive in a way it is not elsewhere.

**A `429` is not a failed run.** Wait `Retry-After` seconds and send the same
request again. Nothing was applied, the device did not move, and the run's step
budget was not spent: a refusal happens before any of that.

A worker that treats any non-2xx as fatal, breaks out of its loop and reports
the run `failed` turns one refusal into one lost run, with whatever actions it
already performed left on the device. That failure mode is the reason this
surface has its own budget, but the budget only widens the gap; it does not
remove it. A worker that does not back off will still find the edge eventually,
and when it does it should slow down rather than give up.

Distinguish the two refusals you can get here, because they call for opposite
responses:

| Status | Meaning | What to do |
| - | - | - |
| `429` | Calling too fast | Wait `Retry-After`, retry the same request |
| `404` | The run is over, or the token is not one | Stop. Do not retry |

A `404` is deliberately the same answer for an expired run, a finished run and a
token that never existed, so it is never worth retrying to find out which.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.