Skip to main content
A run is one autonomous attempt against one device. The control plane owns its state machine. A worker reports what happened; it does not decide what the run’s state becomes. That single sentence is the design. Everything below follows from it.

The lifecycle

There are seven states, and two of them are stops rather than ends.

Goals that need a person

A phonebase-driven run does not carry out goals that move money (pay, buy, order, subscribe, renew, send money), erase an account or everything on the phone, change a password or security setting, close an account, power off or restart the phone, sign out of or remove an account, or sideload an app (“Install unknown apps”). Such a run is created already in needs_user_control with the detail “This goal needs you at the phone. Open the phone and finish it yourself.”, and the takeover bell rings. Resume ends it as succeeded (“You finished this yourself”); it is never replayed. While it runs, the agent also stops with needs_user_control when a screen asks for a payment, a verification code, an account or security setting, or a delete, erase or reset confirmation. Two checks do not depend on the model. On a phone that publishes a UI tree (an emulator or a USB or adb phone), the agent itself reads a fresh tree before every tap, long press, swipe, key press, text entry or app launch, whatever the model looked at, and if the screen is Reset options, Factory data reset, Erase all data, the power menu, a restart prompt, Install unknown apps, the package installer or Remove account (English or Chinese system language), the run stops in needs_user_control with a detail starting risky_screen:<name> and the gesture is not sent. Cloud phones and robot-hand rigs publish no UI tree, so there the agent cannot read the screen itself and only the instructions to the model apply; the power key is withheld from every run instead, and a call to it is refused. If the tree cannot be read at that moment, the gesture goes ahead. Self-hosted runs are unaffected. needs_user_control and paused are easy to confuse, and they are opposites. The first is the agent asking a person for help. The second is a person telling the agent to stop, on a run that asked for nothing. Neither is how a person acts on the device: see Driving alongside a run. Only one of them freezes the run’s deadline: see Putting a run on hold. A run leaves queued for running when the first frame is served to its worker, which is also its first metered step. A run whose deadline passes ends as failed, not as canceled: canceled means somebody chose to stop it. Starting is not the same as working, and the states say which is which. The three terminal states are immutable. Nothing reopens a finished run, and calling cancel_run on one returns it unchanged rather than pretending to act. Terminal is terminal, which is the same discipline the receipt ledger follows. Reaching a terminal state destroys the run’s credential and frees your org’s concurrency slot, so another run can start. It does not release the device lease. That lease is still yours and still accruing active time until you release it or it expires. Release it yourself when the run is done.

Why a run ended

A run that did not simply succeed carries a reason from a closed set: Three of those deserve a note. deadline_exceeded is reached by two very different roads, and detail tells you which. Either no worker ever connected, which is not a timeout at all but a missing participant, or a worker connected and did not finish in time. The structured form of the same fact is workerFirstContactAt: null means nobody came. If it is null, start at Bring your own worker; if it is a timestamp, the reason your worker stopped is in your worker’s own logs. A run parked as needs_user_control can also end this way. Its clock keeps running while it waits for a person and is capped by the lease, so the person has to finish and press Resume inside the run’s remaining time. When that is what happened, detail says so and keeps the instruction the person was given. lease_lost is terminal on purpose. If the lease has gone, every subsequent action would fail the same way, so the run ends immediately rather than leaving a worker spinning against refusals until its deadline, holding one of your concurrency slots for nothing. budget_exhausted refuses the next frame, and does not by itself end the run. A worker that spends its last step finishing the job and then reports success has succeeded. The distinction matters: treating a finished run as a failure would feed it into retry paths and produce a second set of physical actions and a second charge for work that was already done. Deadlines are reclaimed lazily. A run past its deadline is moved to failed the next time anything looks at it, rather than by a background sweeper. That matters for a reason you can feel: an org’s concurrency allowance counts only runs that are neither terminal nor expired, so a crashed worker does not permanently occupy a slot.

Watching a run

get_run is a bounded long poll, not a stream.
  • With just a runId, it answers immediately with the current state.
  • With waitMs and sinceState, it returns as soon as the state differs from sinceState, and answers at the deadline regardless.
It never hangs, and both faces give the same answer to an over-long wait: a waitMs above the maximum is refused, not silently shortened, by the get_run tool and by GET /v1/runs/{runId} alike. Loop on the call rather than trying to hold one open. The maximum is published in the tool reference. What comes back is the run’s public view:
stepSeq is how many metered steps have been issued so far, one per frame served. It is the live view of what this run is costing you.
get_run never echoes the run URL. That credential is returned once, by start_run, and nowhere else. If you lost it, cancel the run and start another rather than looking for a way to read it back.
A run has two ids, and one run. runId is the handle whoever started it kept; agentRunId is the id every receipt the run wrote carries and the one a permalink is built on. There is no retry mechanism here and no parent task: nothing groups several agentRunIds under one runId. Starting again is starting a new run, with both ids new. runId values begin task_ for historical reasons; treat both ids as opaque.

Ending a run

cancel_run ends a run and destroys its credential. An already terminal run is returned unchanged.

Driving alongside a run

A person does not have to stop a run to act on its device. The lease a run holds is your organisation’s, and inside the organisation it is shared: a signed-in person may send a gesture through POST /v1/devices/{deviceId}/actions without naming a lease, and it runs under the run’s lease. The run keeps going. Three rules make that safe rather than chaotic:
  1. One device, one lane. Gestures on a device run one at a time in the order they arrived, whoever sent them, so a swipe is never interleaved with a tap and a receipt is written for each one saying who took it.
  2. A person goes first. A person’s gesture waits for the action that is running to finish and then goes ahead of every action the run has waiting. Nothing the run already started is interrupted.
  3. The agent is told, and by default refused. Every action result and every screenshot the run receives carries others: how many actions principals other than the run took on the device since the frame it last decided on, when the newest was, and whether they were a person or another agent. An action bound to observedFrame is refused stale_frame (reason others_acted) when that count is above zero, before anything physical happens and without consuming your idempotency key. A worker that already re-observes on a stale frame needs no new code; send refuseIfOthersActed: false to be informed instead of stopped.
The person’s own results carry the same others about the agent, plus coDriving: true to say the lease was somebody else’s. Nothing about the lease itself changed: a device still holds one lease at a time across every organisation, the fencing token still fences a stale holder, and the meter still runs on the lease. GET /v1/devices/{deviceId}/driving answers who is driving in one read: the run on the device, whose lease it is, and who else acted in the last sixty seconds.

Letting a parked run continue

When the agent meets something only a person can do, a sign in or a challenge, its worker reports needs_user_control. The run parks rather than ending: it is asking a person to do the next step on the same screen. Parked means the agent is genuinely stopped. No new frames are served to the worker and no actions from it are accepted. Both gates are shut, not only the one that looks at the screen. The person does the step on the device, through the console or the action route, and then lets the run continue:
resume_run lifts either kind of stop, and only a person can lift either. A finished run stays finished, and a run in any other state comes back unchanged rather than being nudged sideways. Resuming is on the human side only. A worker cannot lift its own request, because a worker that could decide the person was finished would make the request meaningless. Secrets never travel through the run either; the person acts on the device. See Progress and takeover.

Putting a run on hold

needs_user_control is the agent asking for a person. This is the reverse: a person telling the agent to stop, on a run that asked for nothing. You do not need it to act on the device; you need it when you want the agent’s hands off the phone while you work.
or, over MCP:
It works from queued, running and needs_user_control alike, and it hands you three things:
leaseId is the lease the run acts under, kept in this response for compatibility. You do not need it to act: a signed-in person’s gestures resolve the lease themselves. It is still a capability, not a detail, which is why GET /v1/leases deliberately never publishes it. Four things follow from putting a run on hold:
  1. The agent’s channels shut. No frames are served to its worker and no actions from it are accepted until you resume.
  2. The run’s deadline freezes. deadlineAt stops advancing while the run is paused, and resuming slides it forward by however long the hold lasted. Time the agent spends on hold is not charged to the run.
  3. The worker’s steps are untouched. Resuming carries on from the current screen, whatever you did on it meanwhile, rather than redoing what is already done.
  4. GET /v1/runs/{runId} reports pausedAt, so a client can tell a run on hold from a run that is merely between steps. Calling pause again on a run already on hold is safe: nothing is re-stamped.
There are two ways the hold ends badly, and they differ: No new failure reason is used for either case. To keep the hold longer, renew the lease. To let the agent continue, resume the run.

Doing it from the console

The console’s /devices/:id Control block is where all of this lives. The screen is always interactive: a tap goes to the phone under whatever lease your organisation holds, and the run, if there is one, keeps going. Pause agent and Stop run are the two interruptions, and both are yours to choose.

The same lifecycle over HTTP

If you are not speaking MCP, the same operations are HTTP routes on the control surface, authenticated with the same bearer credential: GET /v1/runs/{runId} accepts the same bounded wait as a query parameter.

Two ids, one route

GET /v1/runs/{runId} takes either id. Send the runId the start response returned, or the agentRunId a receipt carries and a permalink is built on: an id beginning run_ is read as an agentRunId, anything else as a runId. So a run id somebody shared with you can be read directly instead of by paging the listing until it appears, and both spellings answer with the same run view and the same bounded wait. The four write routes take a runId only. An agentRunId there is 404, the same 404 an unknown id gets. That read returns the run head only. The action trail is a separate, paged resource at GET /v1/receipts?agentRunId=..., and it is not folded in: it has its own cursor and its own retention, and embedding it would mean either truncating it at an unchosen number or running a second pagination contract inside a single resource response. The frames a run saw are archived too, and have no read endpoint today. See Frame retention. Every one of these routes applies the calling key’s device narrowing. A run on a device your key cannot address is 404, the same 404 an unknown id gets. Both faces run the same operation underneath, so there is no second path to diverge. They also share the scope rule: runs:start and devices:act, enforced identically, so a credential that cannot start a run over MCP cannot start one over HTTP either. A run belonging to another org reports as not found (run_not_found). There is no way to learn that a run id exists but is not yours.

What the worker sees

Almost nothing, and that is the point. See The contract in full for the surface a worker actually talks to, and what it is structurally prevented from reaching.