Skip to main content
Work in this order: recognise the shape of the failure, then read its code, then act. Guessing from the message text is the slow path.

First, know which shape you have

An error result arrives with isError: true and comes in exactly two shapes.
  1. A JSON envelope in the text block, carrying a stable code, a message, and retryable. This is a device or control plane error.
  2. Prose in the text block, which will not parse as JSON. This is a schema rejection: your parameters were wrong.
So: try to parse the text block. If it parses, branch on code. If it does not, fix your arguments and do not resend the same call. Errors never carry structuredContent. Every code is listed in the error reference.

Signing up

I submitted the sign-up form and nothing happened. Sign-up runs a short human check in the background. A browser slows down background work in a tab that is not in front, so if you switched to another window while it ran, the check may not have finished and the form waits without saying so. Bring the tab to the front, wait a few seconds, and submit again. What you typed is still in the form.

Connection and authentication

A tool I expect is missing from listTools. That is authorization, not a fault. Scopes are enforced when tools are registered, so a tool outside your credential’s scopes is never registered and genuinely does not exist on your connection. Check your key’s scopes. If the missing tools are the run tools, note they need runs:start and devices:act. See Authentication. 401, with a reason in the body. Read the reason. missing_credentials means no Authorization header arrived. malformed_credentials means the header was not a usable bearer value. invalid_token, token_expired and wrong_audience are about a JWT. unknown_key, revoked_key and expired_key are about an API key. unknown_key deliberately covers both “no such key” and “wrong secret”. There is no way to tell them apart, and that is on purpose. 503 from the authentication path. The credential store or the authorization server could not be reached. This is fail closed by design: an outage refuses requests rather than letting them through. Retry. My client is refused before it authenticates. Two checks run first. The endpoint accepts only the host it was configured for, and it refuses any request carrying a browser Origin header. A client that sets a browser style origin will be refused no matter what credential it has. See Transports. 405 on /mcp. You sent GET or DELETE. The endpoint is stateless, so there is no stream to resume and no session to end. Post your requests. 429, with rate_limited in the body. You are over your request budget. Note the shape: this is a transport level 429 carrying the HTTP error envelope, not a tool result with isError, because the limiter runs before the MCP endpoint is reached. So a client that only inspects tool errors will see a transport failure here. Wait the number of seconds in Retry-After and retry the same request. The refusal happens before any work, so nothing was half applied. Do not retry in a tight loop; the allowance refills on a clock and retrying faster cannot speed it up. RateLimit-Remaining on your earlier responses would have told you this was coming. See Rate limits.

Leases

lease_required. Your organization does not hold that device right now, and this call could not take it for you. Either you named a leaseId that is gone (naming one is a claim that you hold the device, so it is answered rather than quietly turned into a new take), or your credential has no devices:lease and so may act only on a device your organization already holds, or the lease ended while your action waited its turn behind others. Take the device and try again: a gesture sent with no leaseId takes it for you when your credential may. lease_mismatch. The leaseId you supplied does not match the current lease. Actions, run starts, renewals, and releases refuse it before changing the phone. Refresh the lease view, then use its current id, or omit the id where the operation allows that. Authentication and your organization’s possession are still required; knowing an id alone grants no access. acquire_device fails on a device that looks idle. A device holds one lease at a time, across every org. Somebody else holds it, or a previous run of yours did not release. Leases expire on their own, so waiting works, and reading expiresAt from the holder’s acquire response is better than guessing. A person in your org who wants to act on the device does not need to acquire: the console acts under the lease your org holds. An action answered stale_frame with reason others_acted. Somebody else acted on the device after the frame you decided on, and your action was bound to that frame. This is the default for an API key, a copilot or a run that sends observedFrame, and what refuseIfOthersActed: true asks for. Nothing physical happened and your idempotency key was not consumed: take a new screenshot, decide again, act again. Send refuseIfOthersActed: false to be told through others without being refused. An action answered device_busy. Too many actions were already waiting on that device. Actions run one at a time per device in arrival order, and a bounded number may wait. Nothing you already sent was dropped; retry after retryAfterMs. A run ended with lease_lost. The device moved to another holder while the run was going. The run is ended immediately on purpose, so a worker does not spin against refusals for the rest of its deadline while occupying one of your concurrency slots.

Actions and observations

stale_frame. The frame you passed as observedFrame is no longer the device’s current screen. This is the guard working: the screen changed between your decision and your action. Observe again, decide again, act again. A frame does not age out, so this is never about how long you took. Take as long as you need between observing and acting. The exception is a device whose own resource expires frames, which answers reason expired; a non-null freshnessMs in its capabilities is how you know to expect that. Nothing physical was attempted, and your idempotency key was not consumed, so you can retry the same intent under the same key. capability_unsupported. The device cannot do what you asked. Read its capability block: the key you pressed may not be in its pressKey list, it may have no UI tree, or you may have passed observedFrame to a device whose frames are not referenceable. See Capabilities. backend_not_implemented. The device’s backendKind is a declared kind with no adapter behind it yet. The device row is valid; the mechanism does not exist. See Capabilities for which kinds you can reach today. device_unavailable or device_disconnected. The device is not reachable right now. Both are retryable. get_device will tell you its current status without enumerating the fleet. invalid_argument, or a prose rejection. Your parameters. Unknown keys are rejected rather than dropped, so a misspelling fails here rather than reaching the device with a missing value. Coordinates must be on the 0 to 1000 grid.

Retries and outcomes

I never got a response to an action. Do not retry blind. Call get_receipt with the same idempotencyKey.
  • found: false means the action never began. The same key is safe to retry.
  • A committed receipt means it happened. Do not do it again.
  • A failed receipt tells you why it failed.
idempotency_conflict. You reused a key for a different action. A key identifies one intended action, not a session. Use a new key. A receipt says outcome_unknown with requiresManualReview. The transport to the device was lost after the action was dispatched. Whether the device acted is genuinely unknown, and the control plane refuses to guess in either direction. A person has to check the device. A receipt says interrupted_by_restart. The control plane restarted while the action was in flight, and the request was closed as failed rather than left pending. Same treatment: check before repeating anything that matters. action_outcome_unknown. The live form of the same situation. Treat it as “check, do not repeat”.

Agent tier

run_not_found. No run under that id for your org. A run belonging to another org reports identically, so this never tells you a foreign run exists. Also remember that run records live in the control plane process today: a restart loses in flight runs, and a lost run reports as not found. My worker got 409 with “no further steps”. The run reached its step budget, or its deadline passed. The step budget refuses the next frame; it does not by itself end the run. Raise maxSteps on the next run, within the deployment’s own ceiling, which a caller cannot raise. My worker got 503 on screenshot. A frame is metered before it is served, and a frame that cannot be counted is not served. Treat it as retryable, and specifically do not read it as the device being unreachable. See Device surface codes. A run is stuck in needs_user_control. It is parked, waiting for a person, and it stays parked until one acts. No frames are served and no actions accepted meanwhile. A worker cannot lift its own handoff: resume it with resume_run or the equivalent REST call. See Runs. An org cannot start any more runs. There is a concurrency ceiling per org. Runs that are terminal or past their deadline do not count against it, and a run past its deadline is reclaimed the next time anything looks at it, so monitoring a stale run can free the slot.

Local setup

Run doctor before anything else.
It checks the setup problems people actually hit: an old Node version, device rows that do not parse, a control plane that is not answering, and missing Android tooling. Each failing check prints its own fix, and any failure exits non zero. The control plane refuses to start, saying the receipt ledger is held. Another control plane process is already running over the same devices. Two of them would each keep their own view of who holds a lease and could hand one physical device to two callers, so the second refuses at boot. Stop the first, or give this one its own ledger path. I added a device row and my client cannot see it. An MCP client that spawned the control plane over stdio read the rows once, at startup. Restart the client. The HTTP control plane rereads rows on every request. PHONEBASE_DEVICE_ROWS is set and my file edits do nothing. The environment variable takes precedence over the rows file. The CLI warns when it does. Unset it, or edit the value it points at.

Still stuck

Write to support@phonebase.co. A person reads it. Include what you called, what came back, and the time. If the failure was an HTTP response, the fly-request-id header on it lets us find the request itself. If a generated reference page disagrees with a page written in prose, the generated one is right. Everything else, including access and rates, is on phonebase.co.