Skip to main content
There are two ways to use phonebase. Drive it yourself. Your agent calls tools, one gesture at a time, and decides everything. This is deterministic, auditable, and the right choice whenever you know the sequence. Hand over a goal. You describe what you want in a sentence, and an executor observes and acts on your leased device until it gets there or gives up. This is the agent tier. The low level tools stay available the whole time. The agent tier is an addition, not a replacement, and nothing about the surface you already use changes because you started a run.

Who drives a run

Every run has an executor, and there are two. Pass executor on start_run to choose. Leave it out and the default follows your credential: self_hosted for an API key, phonebase for a signed-in person. The response always says which one the run got, so read executor there before relying on it.
A self_hosted run does nothing on its own. Nothing touches the device until a program of yours picks up the vrxBaseUrl, and a run with no worker pointed at it stays queued until its deadline passes. An API key that leaves executor out gets exactly this kind of run.See Bring your own worker for what a worker has to do and a working loop to start from, or pass "executor": "phonebase" to have the control plane drive the run for you.

How a run is put together

Three parties on a self_hosted run, and it is worth knowing which one holds what. On a phonebase run the control plane plays the worker’s part itself, and no run scoped URL exists. A self-hosted worker is yours. It does not hold your credentials, cannot see your fleet, and never touches the real lease or the fencing token. All it has is one URL that reaches one device for one run, and expires. See Bring your own worker.

Starting a run

Acquire the device first. A run acts under your lease, so you have to hold one, and the control plane refuses to start a run against a lease you do not actually hold rather than issuing a credential for nothing.
workerConnected is always false in this response. On a self_hosted run, nothing happens on the device until your own worker connects to vrxBaseUrl, and no worker can have connected yet, because this response is the first place that url exists. On a phonebase run the response carries no vrxBaseUrl, no worker ever connects, and progress shows as stepSeq on the run. maxSteps and deadlineAt are the budget the run actually got, not the one you asked for: an ask above the deployment’s limit is capped to that limit, and the deadline is also bounded by the lease’s own expiry. Read them here rather than assuming your ask was honoured.
On a self_hosted run, vrxBaseUrl contains a credential and is returned here only. Nothing else ever echoes it back, including the tool you use to watch the run. Hand it to your worker and to nothing else: never a log, an issue, a support thread, or a commit.
Keep runId. That is the handle for everything afterwards, and it is safe to pass around. The same operation is available over HTTP for callers that are not speaking MCP, and both faces run the same operation underneath, so they cannot behave differently. See Runs and the API overview.

Budgets live on the control plane

A run carries hard ceilings, and they are enforced where they can actually be enforced.
  • Steps. How many frames the run may consume. maxSteps on your request can only lower it; the deployment’s own ceiling cannot be raised by a caller. Default: 50 steps per run.
  • Wall clock. An absolute deadline, also capped by the deployment and by your lease’s own expiry, whichever comes first. A lease expires 10 minutes after it is granted, and renew_lease slides that expiry, so a run that keeps renewing is bounded by the 60 minute lifetime no lease may exceed rather than by the first 10 minutes. Stop renewing and the run ends at the next expiry.
  • Concurrency. How many runs one org may have going at once. Default: 4 runs per org.
Those three defaults are what a deployment ships with, and a deployment can set its own. They are the numbers to plan against: if you need to know the values in force on the deployment you are pointed at, read them off a rejection, which names the limit it enforced. The reasoning is worth understanding, because it explains where the guard rails are. The expensive part of an autonomous run happens inside the worker, where the control plane has no visibility. But the control plane holds the device channel, so it is the only place that can actually stop one. When the control point and the cost point are in different places, the gate belongs where it can genuinely close. There is one refinement that saves real work. Running out of steps refuses the next frame; it does not kill the run. A worker that spends its last step finishing the job and then reports success is reporting success, not failing on a technicality. The terminal state comes from the worker’s report or from the deadline.

Metering is anchored to frames

One metered step is one frame served to the worker. Not one action. A worker that looks once and then performs five taps has consumed one step. A worker that decides it is finished without acting has still consumed the step it used to decide. The step sequence is issued by the control plane, never accepted from the worker, because the party being metered must not hold the pen. And the count is written before the pixels are served: if the count cannot be recorded, the frame is not served. That ordering has a consequence worth stating in full, because it runs one way only. Delivered implies billed. Billed does not imply delivered: a client that drops the connection between the metering write and the response has been counted for a frame it never read, which costs at most one step per dropped connection. Everything that fails before the write costs nothing. See Metering units for the complete list. This is honest about its own precision. A step is a lower bound on the number of model calls a worker made, not an equivalent: a worker may reason several times over one frame. See Metering units.

Watching and ending

get_run is a bounded long poll. With no arguments beyond the id it answers immediately. With a wait and a state to compare against, it returns as soon as the state differs, and answers at the deadline regardless. It never hangs. cancel_run ends a run and destroys its credential. resume_run lifts a run parked for a person. Both are described in Runs.

Scopes

The run tools need runs:start and devices:act. Handing a device to a loop is a distinct grant from driving it yourself, and the run URL drives the device, so a credential that may not act may not start a run either. See Authentication.

What version 1 does not do

Stated plainly, because finding out later is worse.
  • There is no queue across multiple workers. A run is claimed once.
  • The built in executor is not configurable per run. A phonebase run is driven by the control plane with its own model and prompt. To choose your own, run self_hosted with a worker you write. See Bring your own worker.
  • A run is not resumed across a restart. Its record survives, so you can still read what it was trying to do and how it ended, but nothing picks the work back up on the other side: a run interrupted by a restart ends at its deadline like any other run whose worker stopped answering.
  • A lease can be renewed, up to an hour. It expires 10 minutes after it is granted, and renew_lease slides that expiry another 10 minutes without changing the lease id or the fencing token, so a goal that needs longer keeps the same claim by renewing before the expiry passes. A lease may not live more than 60 minutes from when it was acquired; past that, acquire a new one. The 10-minute expiry is what makes a crashed client give the device back: stop renewing and the device frees itself as it always did.
Runs are readable after they finish: GET /v1/runs lists yours, newest first, each entry carrying the goal it was started with and how it ended, so losing a run id no longer loses the run. Treat a run you cannot find the way you would treat a failed one.