> ## Documentation Index
> Fetch the complete documentation index at: https://docs.phoneuse.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Bring your own worker

> You supply the worker that executes a run. Here is what it has to do and how to get one running.

**You provide the worker.** phonebase does not execute your runs.

`start_run` creates the run, holds its state machine, meters it, and hands you a
`vrxBaseUrl`. Nothing then drives the device until a program of yours picks that
url up and starts asking for frames. If you never point a worker at it, the run
sits in `queued` until its deadline passes and it ends as `failed`.

This page exists because that is easy to miss. Read it before you write your
first `start_run` call.

## Who does what

| Party | Supplied by | Responsibility |
| - | - | - |
| The control plane | phonebase | Run state, lease, fencing token, budgets, metering |
| The device | phonebase | The physical or cloud device your lease points at |
| The worker | **You** | Deciding what to do next, and doing it through VRX |

The split is deliberate. The control plane is hands and eyes. The judgment about
what to tap next is yours, which is what lets you choose your own model, your
own prompt, and your own stopping rule, and change any of them without waiting
for us.

## If you do not want to write one

You do not have to use the agent tier at all. The low level tools drive a device
directly and need no worker:

1. `acquire_device` to take a lease.
2. `screenshot` to see the screen.
3. `tap`, `swipe`, `type_text` to act.
4. `release_device` when you are done.

Every action returns a receipt. This is the shorter path to a first working
action, and it is the right one whenever you already know the sequence of steps.
See [Your first session](/getting-started/first-session).

Use the agent tier when you want a goal executed by a loop instead of a script.

## What a worker has to do

A worker repeats four steps until it is finished:

1. Ask for a frame.
2. Decide what to do, usually by sending the frame to a vision capable model.
3. Send one action.
4. Report the outcome when it is done or stuck.

That is the whole job. A worker holds exactly one credential, the `vrxBaseUrl`,
and it needs nothing else: no API key, no lease id, no headers. If your worker
wants any of those, it has drifted outside the contract.

## The contract in full

Everything a worker may call, rooted at the `vrxBaseUrl` you were handed:

| Method | Path | Purpose |
| - | - | - |
| `GET` | `/devices` | Your run's device. Read `devices[0].id` |
| `GET` | `/devices/{deviceId}/screenshot` | Raw `image/png` or `image/jpeg` bytes |
| `POST` | `/devices/{deviceId}/actions` | One action |
| `POST` | `/report` | The run's outcome |

Because `GET /devices` tells your worker which device it is bound to, the
`vrxBaseUrl` really is the only input it needs.

Action payloads, exactly as the channel accepts them:

```json theme={null}
{ "action": "tap", "x": 540, "y": 1200, "coordinate_space": "screenshot_pixels", "screen": { "width": 1080, "height": 2400 } }
{ "action": "swipe", "x1": 540, "y1": 1800, "x2": 540, "y2": 600, "duration_ms": 300, "coordinate_space": "screenshot_pixels", "screen": { "width": 1080, "height": 2400 } }
{ "action": "type_text", "text": "hello" }
{ "action": "press_button", "button": "home" }
{ "action": "open_app", "package": "com.example.app" }
```

Three rules that catch people out:

* **`screen` must match the frame you just received.** It is checked against the
  current frame, not trusted. Measure the image bytes you were served rather
  than assuming a device resolution, and re-measure after every frame. A
  mismatch is refused with a `400` telling you to re-observe.
* **Coordinates are integers in that frame's own pixels**, and
  `coordinate_space` must be the string `screenshot_pixels`. Taps and swipes
  need both fields. `type_text`, `press_button` and `open_app` do not.
* **Take a frame before your first action.** Acting without ever having observed
  is refused. There is no coordinate basis yet, so there is nothing to bind the
  action to.
* **`type_text` takes at most 4096 characters**, the same ceiling the MCP
  `type_text` tool publishes. Longer text is refused with a `400`, not cut.
* **`duration_ms` on a swipe is a positive integer.** Zero, a fraction or a
  string is refused rather than read as the default. A value above the
  contract's ceiling is performed at the ceiling, because the upstream driver
  sends very large values for a drag.

`press_button` accepts `home`, `back`, `enter`, `delete`, `app_switch`, `power`,
`volume_up` and `volume_down`. Not every device does all of them: the
`capabilities.press_button` field on `GET /devices` tells you which ones this
device advertises, and one it does not support is refused with a `409`.

Report exactly one of three states when you stop:

```json theme={null}
{ "state": "succeeded", "detail": "airplane mode is on" }
{ "state": "failed", "detail": "could not find the settings icon" }
{ "state": "needs_user_control", "detail": "enter the code sent by SMS" }
```

## The loop, in about thirty lines

This is the shape, not a product. Fill in `decide` with a call to the vision
capable model of your choice, and read [Device surface
codes](/mcp/errors#device-surface-codes) for the refusals you have to handle.

```js theme={null}
const base = process.env.VRX_BASE_URL; // from start_run, and a credential

const { devices } = await (await fetch(`${base}/devices`)).json();
const id = devices[0].id;

for (let step = 0; step < 20; step++) {
  const res = await fetch(`${base}/devices/${id}/screenshot`);
  if (res.status === 409) break;               // no more frames, stop and report
  if (!res.ok) continue;                       // 503 is retryable
  const frame = new Uint8Array(await res.arrayBuffer());

  const { width, height } = measure(frame);    // read it from the bytes
  const next = await decide(frame, goal);      // your model call
  if (next.done) break;

  await fetch(`${base}/devices/${id}/actions`, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ ...next.action, coordinate_space: "screenshot_pixels", screen: { width, height } }),
  });
}

await fetch(`${base}/report`, {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({ state: "succeeded", detail: "done" }),
});
```

<Warning>
  The `vrxBaseUrl` carries a credential in its path. Pass it to your worker as an
  environment variable or a secret. Never put it in a log line, a commit, an
  issue, or a support thread.
</Warning>

## Choosing a model

Any model that reads an image and returns a coordinate will work. The quality of
a run is mostly the quality of this decision, so start with the strongest vision
capable model you have access to and reduce only if you measure that a smaller
one holds up.

Your worker pays for its own model calls. phonebase meters frames, not tokens.
One frame is one step no matter how many times you reason over it, so a worker
that thinks twice about the same frame costs you model tokens and costs nothing
extra here. See [Metering units](/billing/units).

## How to tell your worker never connected

The run tells you. `GET /v1/runs/{runId}` and `get_run` both return
`workerFirstContactAt`:

* `null` means no worker has reached this run. A `queued` run with `null` here is
  waiting for you, not for us.
* A timestamp means a worker arrived, and `stepSeq` tells you how many frames it
  has taken since.

If the run ends as `failed` with reason `deadline_exceeded`, `detail` says which
of the two happened: nobody connected, or a worker connected and did not finish.
See [Runs](/agent/runs#why-a-run-ended).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.