> ## Documentation Index
> Fetch the complete documentation index at: https://docs.phoneuse.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat with the copilot

> The chat protocol behind the console's copilot: threads on the server, a streamed UI message protocol, approvals that ride the stream, and the limits a turn runs under.

The copilot is the console's main way to drive a phone: you type what you want,
and it takes the same tools an agent takes, as you, while you watch. This page
is for somebody building a chat client against it, or reading the console's
transport to see what it depends on.

Everything here is served by one route, `POST /v1/copilot`, and four routes for
the threads it writes. All five need a signed in person: a copilot acts as
somebody, and an API key has nobody behind it.

## Threads live on the server

A conversation is a thread, owned by the person who opened it, inside their
organisation. The server keeps it: every message you send, every message the
copilot writes, every tool it called and what came back, and the screenshots
it took along the way.

Three things follow from that.

* **The copilot remembers.** Each turn is shown the thread so far, within a
  budget (see [Limits](#limits)), so you can say "now open the second one"
  without repeating yourself.
* **Your client's history is not trusted.** A request carries the client's
  copy of the transcript, and only its last message is read. What the copilot
  is shown is the stored thread. Two browsers open on one thread see the same
  past, and nobody can put words in the copilot's mouth by editing a payload.
* **A thread is private.** Another person in your organisation cannot list it,
  open it, rename it or delete it. To them it answers 404, the same as a
  thread that never existed. There is no sharing yet.

The four thread routes:

| Route | What it does |
| - | - |
| `GET /v1/copilot/threads` | Your threads, newest activity first, titles and message counts only. `q` searches titles; `limit` and `before` page. |
| `GET /v1/copilot/threads/{id}` | One thread with a page of its messages, newest last, as UI messages. `limit` and `before` page back; `files=none` leaves screenshots out. |
| `PATCH /v1/copilot/threads/{id}` | Rename it. The title is the one field you edit. |
| `DELETE /v1/copilot/threads/{id}` | Delete it. It leaves every listing and read; the id cannot be reused. |

A new thread takes its title from the first message you send, cut to the
length `THREAD_TITLE_MAX_CHARS` allows.

### Reopening a thread

`GET /v1/copilot/threads/{id}` answers with the thread and its newest messages,
oldest first, so the page ends where the conversation stands. It is paged
because a thread is heavy: every tool result carries its full text and every
screenshot its bytes.
The response also carries `mode`: the mode saved on this thread, or `null` if
none was chosen. A second browser can show the same choice when it reopens the
thread.

* `limit` is how many messages, `DEFAULT_THREAD_MESSAGE_ROWS` when omitted and
  at most `MAX_THREAD_MESSAGE_ROWS`.
* `nextBefore` in the response is the oldest message on the page, when older
  ones exist. Send it back as `before` to get the page before it.
* Inline screenshots are carried within `MAX_INLINE_PAGE_BYTES` of decoded
  image data per response (about a third more on the wire, as base64 inside
  JSON). The fit is greedy, newest first: each screenshot is kept if it fits
  what is left, so a large one that does not fit is left out while a smaller
  older one may still be carried. `files=none` leaves every one out. Either
  way `omittedFiles` says how many were left out, and the `frameId` on the
  tool output beside each one still names the picture for the frames API.

Hand `messages` to your chat state as they are. A tool part in state
`approval-requested` on the last message is a question still waiting.

## The protocol

`POST /v1/copilot` answers in the UI message stream when your request lists
`application/vnd.phonebase.ui-message-stream` in `Accept`. The body is the one
the AI SDK's default chat transport sends:

```json theme={null}
{
  "id": "thread_8f2c",
  "trigger": "submit-message",
  "messages": [
    { "id": "u_1", "role": "user", "parts": [{ "type": "text", "text": "Open the camera on pixel-1" }] }
  ]
}
```

`id` is the thread. A new id opens a new thread owned by you. `messages` is
your transcript; only the last message is read, and only its text parts.
User file attachments are rejected with `invalid_argument`. The copilot still
returns screenshots it observes on the phone.

## Which phone the turn is about

A person on the page of one phone who types "open Settings" has already named
the phone. Send what you know beside the message and the copilot will not ask
them to say it again:

```json theme={null}
{
  "id": "thread_8f2c",
  "messages": [{ "id": "u_1", "role": "user", "parts": [{ "type": "text", "text": "Open the camera" }] }],
  "context": {
    "route": "/devices/pixel-1",
    "device": { "id": "pixel-1" },
    "pinned": { "deviceId": "pixel-1" }
  },
  "mentions": ["pixel-2"]
}
```

Both `context` and `mentions` are optional, and both go on either body shape.
`mentions` carries device **ids**, not the names in the sentence, so renaming a
phone can never re-point a message somebody already sent.

The copilot picks the phone in this order and stops at the first hit:

1. `context.pinned.deviceId`, a phone the person pinned to this chat.
2. `mentions`, when it names exactly one.
3. `context.device.id`, the page they are on.
4. The device this thread last acted on, read from its own tool calls.
5. The only device you have, if you have one.
6. Otherwise a `data-device-choice` card.

Three things follow.

* **The copilot is told, and it acts.** A tool call that needs a device and
  names none is given the one the context decided, so a sentence with no phone
  in it still reaches the right phone, and the call you see carries the id.
* **`context` is never replayed.** It describes where somebody is *now*, so
  only this request's copy is read. The stored thread keeps none of it. What
  survives between turns is which device the thread actually touched, which is
  a fact rather than a claim about a page.
* **An id you cannot address is a 400**, naming the field. Every field that is
  wrong is named in the one answer, so you never fix them one round trip at a
  time. `context.device` is `{"id": ...}`, an object; a bare string is refused
  rather than ignored.

`context.device` and the rest may carry more than the id (114 draws them with
a name, a tier and an online flag), and extra keys are accepted. Only the ids
are read. The name, tier and online state the copilot works from are this
service's own, so a page that has gone stale cannot tell it a phone is up.

The response is `text/event-stream`, with the header
`x-vercel-ai-ui-message-stream: v1` so a transport can tell it apart from the
older event stream without reading the body. Each frame is `data:` followed
by one JSON chunk, and the stream ends with `data: [DONE]`.

The chunks, in the order a turn produces them:

| Chunk | Meaning |
| - | - |
| `start` | First. `messageId` names the assistant message this stream writes, which is the message a later approval answer continues. `messageMetadata.createdAt` is when it began. |
| `data-allowance` | `{percent, metered}`: how much of your organisation's copilot allowance is spent. Sent after `start` and again before `finish`. Transient: it reaches your data callback and never becomes a part of the message, so a replayed thread carries none. When `metered` is false, nothing on this deployment is counting yet, and `percent` is not a measurement. |
| `start-step`, `finish-step` | One model step. A step is one answer from the model, with the tools it asked for. |
| `text-start`, `text-delta`, `text-end` | The copilot's prose, as it is written. |
| `reasoning-start`, `reasoning-delta`, `reasoning-end` | The model's reasoning, when the model provides it. |
| `tool-input-start`, `tool-input-delta`, `tool-input-available` | A tool call is being formed, then its arguments are complete. Nothing has run yet. |
| `tool-approval-request` | The call needs your yes. See [Approvals](#approvals). |
| `tool-output-available` | The tool finished. `output` is `{summary, text, frameId?}`: a bounded line, the tool's full answer, and the frame id when a screenshot produced one. |
| `tool-output-error` | The tool's own verdict was an error, or nobody answered an approval in time. `errorText` says which, and is always present because the AI SDK requires it. When the tool itself failed, four more fields ride beside it: `message`, the error's own sentence; `code`, its code (`tool_failed` when the tool's answer was not a readable error); `retryable`, whether the same call may succeed if sent again unchanged; and `retryAfterMs`, how long to wait first, only when the tool said (for example `device_busy`). They are on the stream only: the SDK's `useChat` does not copy them onto the message part, and a replayed thread does not carry them, so read them in your transport if you need them. |
| `tool-output-denied` | You said no. |
| `file` | A screenshot, as a `data:` URL with its media type, when it is small enough to carry inline (`MAX_INLINE_FILE_BYTES`). Larger ones send the `frameId` on the tool output instead, for the frames API. |
| `data-device-choice` | More than one phone could have been meant and none was named. `{reason, devices: [{id, label}]}`: `reason` is the tool that needed a phone. Show the list; the person's pick comes back as the next message, or as `context.device`. The call that raised it is answered `tool-output-error`, and nothing touched a phone. Unlike `data-allowance` this **is** a part, so it is still there when the thread is reopened. |
| `data-suggestions` | Up to three follow-ups, `{items: [{text, deviceId}]}`, each naming a device. Sent before `finish` on a turn that completed. Fill them into the composer; do not send them. Also a part. |
| `finish` | Last. `finishReason` is `stop`, `tool-calls` (an approval is waiting), `length` (the step budget), or `other`. `messageMetadata.stopReason` is the copilot's own closed set, finer than that. |
| `error` | The turn ended badly. Nothing follows it, `finish` included. |

Between chunks the stream may carry SSE comment lines as keep-alives. An
event-source parser ignores them.

Closing the request stops the turn. The copilot finishes the tool call it is
on, if one is running, and stops before its next step or tool call. What it
wrote so far is kept, and the stored message ends with `finishReason: "stop"`.

## Approvals

Two things ask you first: **installing an app**, and **starting an agent run**
that goes on acting by itself after the turn ends. Everything else, looking,
taking and letting go of a device, and every gesture, simply happens because
you asked for it. A card in front of every tap is not consent; it is a
clickthrough, and it teaches people to approve without reading.

The question rides the stream and ends the turn:

1. The copilot sends `tool-input-available` for the call, then
   `tool-approval-request` with an `approvalId` and a descriptor: the tool,
   its arguments, one sentence to show, and `expiresAt`. The sentence is built
   from the tool and its arguments, never from anything the model wrote, so a
   card cannot describe a tap as a screenshot.
2. The stream ends with `finish{finishReason: "tool-calls"}`. The stored
   assistant message holds the pending question, so the answer may arrive on
   any machine, minutes later.
3. You answer by sending the same thread again, with that assistant message
   last and its tool part in state `approval-responded`, carrying
   `approval: {id, approved, reason?, input?, descriptor?}`. This is what the
   AI SDK's `addToolApprovalResponse` produces, plus the two fields below.
4. The next stream continues the same assistant message: its `start` names the
   same `messageId`. On a yes the copilot looks at the screen first, then
   runs the call, then the calls the model had queued behind it, then decides
   again. On a no it sends `tool-output-denied` and the turn ends: a copilot
   that carried on after a refusal, working around it, would make its
   approvals meaningless.

### What the card says

The descriptor decides which card to draw and why, so the client never has to
work it out from the tool name:

| Field | Meaning |
| - | - |
| `kind` | `light` for one line and one row, `plan` for a goal with a budget. A run gets the plan card; it is the one call whose effects outlive the conversation. |
| `reason.code` | `start_run`, `install_app`, `dangerous_operation` or `boundary`. What a receipt records and what you branch on. |
| `summary` | A line under the prompt, in the product's words. Absent when the call declares no ceiling; it never claims one. |
| `budget` | `{maxSteps?, deadlineMs?}`: what the plan card shows and lets somebody edit. |
| `calls` | Present only when one question covers several phones: `[{toolCallId, tool, input}]`. |
| `expiresAt` | Only on the older wire. Absent here, and an absent value never expires. |

A descriptor written before these fields existed carries neither `kind` nor
`reason`. Read a missing `kind` as `light`.

### One card, several phones

Ask for the same thing on three phones and you get one question, not three.
The descriptor's `calls` lists every call it covers, one yes runs all of them,
and each still runs as its own call and leaves its own receipt. A dangerous
operation is never folded into that list: it always gets its own card, so a
yes to three of something can never have been a yes to something else standing
beside them.

### Editing before you agree

On a plan card somebody may change the goal, the step ceiling or the time
limit before saying yes. Send the changed fields as `approval.input`.

```json theme={null}
{ "id": "ap_1", "approved": true, "input": { "maxSteps": 4, "goal": "log in and stop" } }
```

It is **merged** over the arguments you were shown, so sending only what
changed does not delete the rest, and it is re-checked against the tool's own
published input schema. An edit the schema refuses is a 400 naming the field,
in the words of the field rather than a schema path, and nothing runs.

The stored tool part keeps `approval.input` after the call has run, so a thread
read back later still shows the edit the call ran with.

If you echo `approval.descriptor.reason` back, it has to be the code the
question was asked with, or the answer is refused: consent shown for one
reason is not consent for another. Echoing nothing is accepted.

A question that nobody answers within `APPROVAL_TIMEOUT_MS` expires: the answer
is refused with `tool-output-error` on the call. Sending a new message while a
question is pending closes the question the same way.

Only the person the question was asked of can answer it, because the thread is
theirs. An answer to a question that is not pending, or that this thread never
asked, is 404.

## How much it does on its own

Every chat runs in one of three modes. The product default is `auto`, and that
is a deliberate choice rather than a convenience: a card in front of every
action is not consent, it is a clickthrough.

| Mode | What it means |
| - | - |
| `auto` | The default. Direct device actions run; each approval-required plan still asks. |
| `ask` | Every action in the ask tier raises a card. |
| `manual` | The copilot looks and explains. It does not act; you do. |

Send `mode` on a chat request to change that chat's mode. It takes effect on
the turn that sets it and is saved with the thread. An approval authorizes
only the pending action. There are no remembered grants or organisation-wide
automation defaults.

### What is never turned off

No mode bypasses any of these.

* A tool that declares `requiresUserInteraction` in its own MCP metadata.
  This is enforced at every call. Historical `install_app` approval cards
  remain readable, but approving one cannot restore retired installation.
* A text message to somebody this chat has not written to before.
* Doing something to several phones in one go.
* Creating a schedule.
* Money, API keys and team membership, which are never done on your behalf at
  all. Those are not cards; they are refusals with a link.

There is no classifier reading what the copilot is trying to do, and no
threshold that stops it after N similar actions. A goal it set itself does not
widen anything, and neither does text it read off a screen.

## Picking a cut stream back up

A phone that went to the background, a tab switched away from, a network that
dropped. Reconnect to `GET /v1/copilot/threads/{id}/stream` and the rest of
the message arrives as the same chunks, under the same header, ending with the
same `data: [DONE]`.

* `since` is the `messageId` from the `start` chunk you were reading.
* `afterPart` is how many of its parts you already hold. `0`, the default,
  replays the whole message; a number past its end sends the frame and nothing
  inside it, which is how you learn you are caught up.
* Omit `since` and the newest assistant message is replayed from the
  beginning, which is what a client that lost its place entirely needs.
* A `since` naming no message of this thread replays nothing rather than
  restarting from the top. Restarting would repeat what you already have and
  hide that your cursor is stale.

No model is asked anything and nothing is written, so this costs no tokens and
answers the same on a deployment with no copilot configured.

**What it does not do.** It does not rejoin a turn that is still running,
because after a disconnect there is not one: closing the request aborts the
turn, the copilot stops before its next step or tool call, and what it had is
stored. A reconnect therefore gives you everything you missed up to the moment
the connection died. The turn does not carry on in the background.

## Commands: one tool, no model

Sometimes there is nothing to work out. `POST /v1/copilot/commands` runs a
single tool you name, as you, with no model involved and no tokens spent. It is
what the console's `/screenshot` and `/press_key home` become.

```json theme={null}
{ "tool": "press_key", "input": { "key": "home" }, "context": { "device": { "id": "pixel-1" } }, "threadId": "thread_8f2c" }
```

It reaches the same tool surface a chat turn drives, under the same permissions
and the same approval table, so a command cannot do what the conversation
would refuse. Two things differ, and both follow from your having typed the
tool yourself.

* **It works with no copilot configured.** A deployment with no model answers
  503 on `POST /v1/copilot` and still runs commands. Pressing Home should not
  need a language model.
* **The receipt is not marked `via copilot`.** You did this.

`context` and `mentions` work exactly as they do on the chat route, so a
command typed on a phone's page needs no `deviceId`.

What you may drive this way: every read, plus `open_app`, `press_key`,
`release_device`, `list_runs`, `cancel_run` and `resume_run`. Those are the
tools whose arguments you can write out in full. `tap`, `swipe`, `type_text`
and `long_press` are not, because the useful form of "tap the login button" is
a sentence and the coordinates are the copilot's job; nor is `start_run`, which
takes a goal and a budget. Asking for one of those here is a 400 telling you to
use the conversation. A tool your session has no scope for is a different
answer, 403, and the message names the scope, because that remedy is an
administrator rather than a rephrasing.

A 200 carries `{tool, isError, output: {summary, text, frameId?}}` and
optionally `files`. `isError` is the tool's own verdict, not the HTTP status: a
device refusing a gesture is a 200 with `isError` true, exactly as
`tool-output-error` is on the stream.

Send `threadId` and the result is appended to that thread as a tool part. That
is what makes a command followed by a sentence work: the next turn sees what
the command did rather than a gap where it happened. Leave it out and the
command still runs, in no conversation.

If a tool asks before it runs, the answer is 202 with an `approval` descriptor
and an `approvalId`, and nothing has run. Show the card, then send the same
`tool` and `input` again with that `approvalId`. The question is stored on the
thread, so it can be answered minutes later and on another machine, which is
why a command that asks needs a `threadId`. The tool and arguments are checked
against what was asked: a yes to one command cannot be spent on another.

## When it fails

A turn that ends badly sends `error` and nothing after it, with a code from a
closed set. The three you will want to tell apart:

| Code | What it means, and what to show |
| - | - |
| `copilot_unfunded` | The copilot has no credit on our side. The message is the one to show: "The copilot is out of credit on our side. Your phone and / commands still work. We've been told." Nothing the person did caused it and nothing they can do fixes it, so do not send them looking. Their phone, their commands and everything outside the chat are unaffected. We are alerted. |
| `rate_limited` | The model provider is throttling. Worth retrying in a moment. |
| `model_unavailable` | The model could not be reached. Different from both of the above, which is why they are no longer the same code. |

`model_refused`, `tool_failed` and `internal` complete the set. The first three
used to be one code, so a client could only ever say the vaguest of the three.

## Limits

* A turn takes at most `MAX_TURN_STEPS` model steps, across the requests an
  approval splits it into. Reaching the budget is not an error: the work done
  is done, and its receipts exist. A job that genuinely needs more is an
  agent run.
* One message carries at most `MAX_COPILOT_PROMPT_CHARS` characters of text.
* The history the model is shown is the most recent messages that fit within
  `HISTORY_BUDGET_CHARS` characters of transcript, cut at a message of yours so
  the transcript never opens with the copilot's reply. Your newest message is
  always included. What the cut removes is **not forgotten**: it is folded into
  a short summary the copilot writes once and keeps with the thread, shown in
  front of the recent messages on every turn after that. A long conversation
  therefore still knows what it was started for and which phone it was about.
  The summary is regenerated only when the cut reaches past what it already
  covers, so it costs one extra model call when a thread first crosses the
  budget and nothing on the turns after. Nothing that can be rebuilt from the
  server goes into it: the mode, what has been remembered and which phone the
  chat is about are regenerated fresh every turn, so a summary of them would
  be a second and stale copy.
* Screenshots travel inline up to `MAX_INLINE_FILE_BYTES`; above that you get
  the frame id.
* The listing returns at most `MAX_THREAD_ROWS` threads per page, a thread
  read at most `MAX_THREAD_MESSAGE_ROWS` messages, and both refuse a larger
  `limit` rather than quietly reducing it.
* The copilot gives up on a model that stays silent for `MODEL_TIMEOUT_MS`
  within one step, and says so as an error of its own; a model that keeps
  answering slowly is never cut off.

The named limits are constants in the control plane's source, and the OpenAPI
document under [API](/api) carries their current values.

## The older event stream

Until the console has switched, the same route also answers requests that do
not ask for the UI message stream, with `{prompt}` in and named SSE events out
(`session.start`, `message.delta`, `tool.call`, `tool.result`,
`approval.request`, `tool.denied`, `ping`, `done`, `error`). Approvals on that
wire are answered on `POST /v1/copilot/approvals` while the stream stays open.
It goes in the release after the switch; build against the protocol above.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.