Skip to main content
You provide the worker. phonebase does not execute your runs. start_run creates the run, holds its state machine, meters it, and hands you a vrxBaseUrl. Nothing then drives the device until a program of yours picks that url up and starts asking for frames. If you never point a worker at it, the run sits in queued until its deadline passes and it ends as failed. This page exists because that is easy to miss. Read it before you write your first start_run call.

Who does what

The split is deliberate. The control plane is hands and eyes. The judgment about what to tap next is yours, which is what lets you choose your own model, your own prompt, and your own stopping rule, and change any of them without waiting for us.

If you do not want to write one

You do not have to use the agent tier at all. The low level tools drive a device directly and need no worker:
  1. acquire_device to take a lease.
  2. screenshot to see the screen.
  3. tap, swipe, type_text to act.
  4. release_device when you are done.
Every action returns a receipt. This is the shorter path to a first working action, and it is the right one whenever you already know the sequence of steps. See Your first session. Use the agent tier when you want a goal executed by a loop instead of a script.

What a worker has to do

A worker repeats four steps until it is finished:
  1. Ask for a frame.
  2. Decide what to do, usually by sending the frame to a vision capable model.
  3. Send one action.
  4. Report the outcome when it is done or stuck.
That is the whole job. A worker holds exactly one credential, the vrxBaseUrl, and it needs nothing else: no API key, no lease id, no headers. If your worker wants any of those, it has drifted outside the contract.

The contract in full

Everything a worker may call, rooted at the vrxBaseUrl you were handed: Because GET /devices tells your worker which device it is bound to, the vrxBaseUrl really is the only input it needs. Action payloads, exactly as the channel accepts them:
Three rules that catch people out:
  • screen must match the frame you just received. It is checked against the current frame, not trusted. Measure the image bytes you were served rather than assuming a device resolution, and re-measure after every frame. A mismatch is refused with a 400 telling you to re-observe.
  • Coordinates are integers in that frame’s own pixels, and coordinate_space must be the string screenshot_pixels. Taps and swipes need both fields. type_text, press_button and open_app do not.
  • Take a frame before your first action. Acting without ever having observed is refused. There is no coordinate basis yet, so there is nothing to bind the action to.
  • type_text takes at most 4096 characters, the same ceiling the MCP type_text tool publishes. Longer text is refused with a 400, not cut.
  • duration_ms on a swipe is a positive integer. Zero, a fraction or a string is refused rather than read as the default. A value above the contract’s ceiling is performed at the ceiling, because the upstream driver sends very large values for a drag.
press_button accepts home, back, enter, delete, app_switch, power, volume_up and volume_down. Not every device does all of them: the capabilities.press_button field on GET /devices tells you which ones this device advertises, and one it does not support is refused with a 409. Report exactly one of three states when you stop:

The loop, in about thirty lines

This is the shape, not a product. Fill in decide with a call to the vision capable model of your choice, and read Device surface codes for the refusals you have to handle.
The vrxBaseUrl carries a credential in its path. Pass it to your worker as an environment variable or a secret. Never put it in a log line, a commit, an issue, or a support thread.

Choosing a model

Any model that reads an image and returns a coordinate will work. The quality of a run is mostly the quality of this decision, so start with the strongest vision capable model you have access to and reduce only if you measure that a smaller one holds up. Your worker pays for its own model calls. phonebase meters frames, not tokens. One frame is one step no matter how many times you reason over it, so a worker that thinks twice about the same frame costs you model tokens and costs nothing extra here. See Metering units.

How to tell your worker never connected

The run tells you. GET /v1/runs/{runId} and get_run both return workerFirstContactAt:
  • null means no worker has reached this run. A queued run with null here is waiting for you, not for us.
  • A timestamp means a worker arrived, and stepSeq tells you how many frames it has taken since.
If the run ends as failed with reason deadline_exceeded, detail says which of the two happened: nobody connected, or a worker connected and did not finish. See Runs.