Skip to main content
Every request to the hosted HTTP surfaces passes a limiter before it reaches a route. The limiter sits ahead of routing and ahead of authentication, so being refused here tells you nothing about whether your credential was any good. One thing is answered before the limiter and never counted: a browser’s CORS preflight. See If you are calling from a browser. The numbers on this page are defaults. Every response that a budget applied to carries that budget in its headers, and those headers are authoritative: read them rather than hard coding what you find here.

The three budgets

Outside the device surface, which budget applies is decided by one thing: whether this request carries an Authorization header. It is a property of the request, not of the endpoint you aimed it at, so a client that adds a header from shared configuration changes its own budget. On the device surface the run wins, and the header is ignored there for pacing exactly as it is ignored for authorization. Two things follow from “counted per credential” that are easy to get wrong:
  • The budget is per credential, not per org. Two API keys in the same org hold two independent budgets, and one key exhausting its own cannot slow the other down. Respelling the header does not buy you a second budget: the credential is identified the same way the authenticator identifies it.
  • A rotated key starts fresh. A new token is a new budget. The limiter keeps only a hash, never the credential itself.
There is also a ceiling on how many distinct credentials one client IP may hold separate budgets for at once. Past it, a credential that address has not been using is counted against its bare budget instead, and a slot is released once its credential goes quiet. Reaching the ceiling takes more than a handful of credentials in simultaneous use from one address, so if that is your shape, say so when you ask for access. It exists so that presenting a credential cannot itself become a way around the bare budget.

If you are calling from a browser

A browser sends a preflight (OPTIONS) before any request that carries an Authorization header. On a deployment that has granted your origin, the preflight is answered before the limiter and is not counted against any budget, yours or the bare one. It carries no RateLimit-* headers, because none applied. On a deployment with no origins granted there is no preflight handling at all, and an OPTIONS is counted like any other bare request. That is deliberate. A preflight never carries a credential, so counting it could only put it in the bare budget, shared with every other caller behind your egress address. And a refused preflight would never reach your code as a 429: browsers treat any non-2xx preflight as a network error, with no status and no headers to read. So the API never refuses a preflight for rate, and a CORS error in your browser console is a CORS problem, not a rate limit. The request that follows the preflight is counted against your credential in the ordinary way. If it is refused, you get a real 429: the response carries Access-Control-Allow-Origin, and Retry-After and the RateLimit-* headers are exposed, so your code can read them and back off. One caveat that applies to browsers more than to anything else: a session token that rotates presents a new credential on every rotation, and the ceiling on distinct credentials per client IP described above is reached by rotation alone. Past it, requests are counted against the bare budget until an older credential goes quiet, so a browser can find itself on the smaller budget without having changed anything it does. For bare requests the client IP is the one our edge recorded for the connection, falling back to the first hop in X-Forwarded-For.

The device surface

The device surface, /vrx/{runToken}/v1/..., carries its credential in the URL path and ignores any Authorization header, by design. It gets its own budget, counted per run. That means two runs behind one egress address do not share an allowance, and one run spending its own cannot slow another down. Adding workers behind the same address adds capacity rather than dividing it. The budget applies to a run the control plane currently holds live. A request naming a run token that has expired, ended, or never existed is counted against the bare budget for your client IP, alongside every other uncredentialed caller behind that address. That is the same answer such a request has always had, so a worker still calling after its run is over pays the bare rate. The switch happens when the run ends, not when you next call, so there is no window in which a finished run is still spending at the run rate.
A run’s budget bounds how fast it may call, not how much it may do in total. The total is the run’s step budget, which is fixed when the run starts and enforced before any frame is issued. Raising one does not raise the other.
Two consequences worth planning around:
  • The budget arrives with the run. It is in force from your first request, so there is nothing to register and nothing to wait for.
  • It does not survive the run. The moment the run reaches a terminal state its token stops working and its budget is withdrawn, so the next call gets a 404 counted against the bare budget.

Being refused

Over budget is 429, with the control plane’s HTTP error envelope:
This is the HTTP envelope, not the MCP tool error envelope. The limiter runs before the MCP endpoint is reached, so a throttled tool call comes back as a transport level 429, not as a result with isError and a code from the error reference. A client that only branches on tool errors will see this as a transport failure, which is what it is. Retry-After on that response is an integer number of seconds, computed from how far under one request the caller actually is. It is not a fixed backoff constant, so honour it rather than replacing it with your own.

The headers

Every response that a budget applied to carries these: Both counts describe the allowance at the moment the request was admitted, not at the moment you read them. On a long running call the real figure has already crept up, so they understate rather than overstate what you have left. RateLimit-Reset and Retry-After answer different questions and will differ. Reset is seconds until full. Retry-After on a 429 is seconds until one request is available. Pacing against Reset when you only need one call will make you wait longer than necessary. The allowance refills continuously rather than resetting on a boundary, so a caller that has been idle can spend up to the full budget at once and then drains toward the steady rate. Absence is information. No RateLimit-* headers on a response means no budget applied to it, either because the path is exempt or because that caller class is switched off in this deployment. It never means “unknown”.

What is exempt

GET and HEAD on /healthz are never limited, in any caller class. A 429 on a health probe would read as an outage and mask the real ones. Only those two methods are exempt, and only that exact path. On a deployment that has granted at least one origin, a CORS preflight on a path a browser may reach (/v1/*, /healthz and the discovery documents) is answered before the limiter and never counted. Only a real preflight qualifies: an OPTIONS request that carries an Origin header. An OPTIONS without one is an ordinary request and is counted like any other.

What the budgets are scoped to

Two boundaries worth knowing before you plan around these numbers.
  • Per serving process. The allowance lives in the process that answers you. Read the numbers as a floor you can rely on rather than a global ceiling on the deployment, and pace against the headers you actually receive.
  • Hosted only. Local mode has no limiter at all. It binds to loopback and serves one operator, so there is nothing to protect and no headers to read. Do not use local behaviour to predict hosted behaviour here.

Living within them

  • Read RateLimit-Remaining and slow down before you are refused. Being refused costs you a round trip and tells you nothing you could not have read from the previous response.
  • On a 429, wait Retry-After seconds and retry the same request. The refusal happens before any work, so nothing was half applied and the retry is safe.
  • Do not retry a 429 immediately or in a tight loop. The allowance refills on a clock, and retrying faster cannot make it refill faster.
  • If your steady state does not fit the default budget, say so when you ask for access. The budget is a deployment setting, not a property of the protocol.

If you are driving a device

One rule matters more on this surface than anywhere else, because getting it wrong is expensive in a way it is not elsewhere. A 429 is not a failed run. Wait Retry-After seconds and send the same request again. Nothing was applied, the device did not move, and the run’s step budget was not spent: a refusal happens before any of that. A worker that treats any non-2xx as fatal, breaks out of its loop and reports the run failed turns one refusal into one lost run, with whatever actions it already performed left on the device. That failure mode is the reason this surface has its own budget, but the budget only widens the gap; it does not remove it. A worker that does not back off will still find the edge eventually, and when it does it should slow down rather than give up. Distinguish the two refusals you can get here, because they call for opposite responses: A 404 is deliberately the same answer for an expired run, a finished run and a token that never existed, so it is never worth retrying to find out which.