Skip to main content
The copilot is the console’s main way to drive a phone: you type what you want, and it takes the same tools an agent takes, as you, while you watch. This page is for somebody building a chat client against it, or reading the console’s transport to see what it depends on. Everything here is served by one route, POST /v1/copilot, and four routes for the threads it writes. All five need a signed in person: a copilot acts as somebody, and an API key has nobody behind it.

Threads live on the server

A conversation is a thread, owned by the person who opened it, inside their organisation. The server keeps it: every message you send, every message the copilot writes, every tool it called and what came back, and the screenshots it took along the way. Three things follow from that.
  • The copilot remembers. Each turn is shown the thread so far, within a budget (see Limits), so you can say “now open the second one” without repeating yourself.
  • Your client’s history is not trusted. A request carries the client’s copy of the transcript, and only its last message is read. What the copilot is shown is the stored thread. Two browsers open on one thread see the same past, and nobody can put words in the copilot’s mouth by editing a payload.
  • A thread is private. Another person in your organisation cannot list it, open it, rename it or delete it. To them it answers 404, the same as a thread that never existed. There is no sharing yet.
The four thread routes: A new thread takes its title from the first message you send, cut to the length THREAD_TITLE_MAX_CHARS allows.

Reopening a thread

GET /v1/copilot/threads/{id} answers with the thread and its newest messages, oldest first, so the page ends where the conversation stands. It is paged because a thread is heavy: every tool result carries its full text and every screenshot its bytes. The response also carries mode: the mode saved on this thread, or null if none was chosen. A second browser can show the same choice when it reopens the thread.
  • limit is how many messages, DEFAULT_THREAD_MESSAGE_ROWS when omitted and at most MAX_THREAD_MESSAGE_ROWS.
  • nextBefore in the response is the oldest message on the page, when older ones exist. Send it back as before to get the page before it.
  • Inline screenshots are carried within MAX_INLINE_PAGE_BYTES of decoded image data per response (about a third more on the wire, as base64 inside JSON). The fit is greedy, newest first: each screenshot is kept if it fits what is left, so a large one that does not fit is left out while a smaller older one may still be carried. files=none leaves every one out. Either way omittedFiles says how many were left out, and the frameId on the tool output beside each one still names the picture for the frames API.
Hand messages to your chat state as they are. A tool part in state approval-requested on the last message is a question still waiting.

The protocol

POST /v1/copilot answers in the UI message stream when your request lists application/vnd.phonebase.ui-message-stream in Accept. The body is the one the AI SDK’s default chat transport sends:
id is the thread. A new id opens a new thread owned by you. messages is your transcript; only the last message is read, and only its text parts. User file attachments are rejected with invalid_argument. The copilot still returns screenshots it observes on the phone.

Which phone the turn is about

A person on the page of one phone who types “open Settings” has already named the phone. Send what you know beside the message and the copilot will not ask them to say it again:
Both context and mentions are optional, and both go on either body shape. mentions carries device ids, not the names in the sentence, so renaming a phone can never re-point a message somebody already sent. The copilot picks the phone in this order and stops at the first hit:
  1. context.pinned.deviceId, a phone the person pinned to this chat.
  2. mentions, when it names exactly one.
  3. context.device.id, the page they are on.
  4. The device this thread last acted on, read from its own tool calls.
  5. The only device you have, if you have one.
  6. Otherwise a data-device-choice card.
Three things follow.
  • The copilot is told, and it acts. A tool call that needs a device and names none is given the one the context decided, so a sentence with no phone in it still reaches the right phone, and the call you see carries the id.
  • context is never replayed. It describes where somebody is now, so only this request’s copy is read. The stored thread keeps none of it. What survives between turns is which device the thread actually touched, which is a fact rather than a claim about a page.
  • An id you cannot address is a 400, naming the field. Every field that is wrong is named in the one answer, so you never fix them one round trip at a time. context.device is {"id": ...}, an object; a bare string is refused rather than ignored.
context.device and the rest may carry more than the id (114 draws them with a name, a tier and an online flag), and extra keys are accepted. Only the ids are read. The name, tier and online state the copilot works from are this service’s own, so a page that has gone stale cannot tell it a phone is up. The response is text/event-stream, with the header x-vercel-ai-ui-message-stream: v1 so a transport can tell it apart from the older event stream without reading the body. Each frame is data: followed by one JSON chunk, and the stream ends with data: [DONE]. The chunks, in the order a turn produces them: Between chunks the stream may carry SSE comment lines as keep-alives. An event-source parser ignores them. Closing the request stops the turn. The copilot finishes the tool call it is on, if one is running, and stops before its next step or tool call. What it wrote so far is kept, and the stored message ends with finishReason: "stop".

Approvals

Two things ask you first: installing an app, and starting an agent run that goes on acting by itself after the turn ends. Everything else, looking, taking and letting go of a device, and every gesture, simply happens because you asked for it. A card in front of every tap is not consent; it is a clickthrough, and it teaches people to approve without reading. The question rides the stream and ends the turn:
  1. The copilot sends tool-input-available for the call, then tool-approval-request with an approvalId and a descriptor: the tool, its arguments, one sentence to show, and expiresAt. The sentence is built from the tool and its arguments, never from anything the model wrote, so a card cannot describe a tap as a screenshot.
  2. The stream ends with finish{finishReason: "tool-calls"}. The stored assistant message holds the pending question, so the answer may arrive on any machine, minutes later.
  3. You answer by sending the same thread again, with that assistant message last and its tool part in state approval-responded, carrying approval: {id, approved, reason?, input?, descriptor?}. This is what the AI SDK’s addToolApprovalResponse produces, plus the two fields below.
  4. The next stream continues the same assistant message: its start names the same messageId. On a yes the copilot looks at the screen first, then runs the call, then the calls the model had queued behind it, then decides again. On a no it sends tool-output-denied and the turn ends: a copilot that carried on after a refusal, working around it, would make its approvals meaningless.

What the card says

The descriptor decides which card to draw and why, so the client never has to work it out from the tool name: A descriptor written before these fields existed carries neither kind nor reason. Read a missing kind as light.

One card, several phones

Ask for the same thing on three phones and you get one question, not three. The descriptor’s calls lists every call it covers, one yes runs all of them, and each still runs as its own call and leaves its own receipt. A dangerous operation is never folded into that list: it always gets its own card, so a yes to three of something can never have been a yes to something else standing beside them.

Editing before you agree

On a plan card somebody may change the goal, the step ceiling or the time limit before saying yes. Send the changed fields as approval.input.
It is merged over the arguments you were shown, so sending only what changed does not delete the rest, and it is re-checked against the tool’s own published input schema. An edit the schema refuses is a 400 naming the field, in the words of the field rather than a schema path, and nothing runs. The stored tool part keeps approval.input after the call has run, so a thread read back later still shows the edit the call ran with. If you echo approval.descriptor.reason back, it has to be the code the question was asked with, or the answer is refused: consent shown for one reason is not consent for another. Echoing nothing is accepted. A question that nobody answers within APPROVAL_TIMEOUT_MS expires: the answer is refused with tool-output-error on the call. Sending a new message while a question is pending closes the question the same way. Only the person the question was asked of can answer it, because the thread is theirs. An answer to a question that is not pending, or that this thread never asked, is 404.

How much it does on its own

Every chat runs in one of three modes. The product default is auto, and that is a deliberate choice rather than a convenience: a card in front of every action is not consent, it is a clickthrough. Send mode on a chat request to change that chat’s mode. It takes effect on the turn that sets it and is saved with the thread. An approval authorizes only the pending action. There are no remembered grants or organisation-wide automation defaults.

What is never turned off

No mode bypasses any of these.
  • A tool that declares requiresUserInteraction in its own MCP metadata. This is enforced at every call. Historical install_app approval cards remain readable, but approving one cannot restore retired installation.
  • A text message to somebody this chat has not written to before.
  • Doing something to several phones in one go.
  • Creating a schedule.
  • Money, API keys and team membership, which are never done on your behalf at all. Those are not cards; they are refusals with a link.
There is no classifier reading what the copilot is trying to do, and no threshold that stops it after N similar actions. A goal it set itself does not widen anything, and neither does text it read off a screen.

Picking a cut stream back up

A phone that went to the background, a tab switched away from, a network that dropped. Reconnect to GET /v1/copilot/threads/{id}/stream and the rest of the message arrives as the same chunks, under the same header, ending with the same data: [DONE].
  • since is the messageId from the start chunk you were reading.
  • afterPart is how many of its parts you already hold. 0, the default, replays the whole message; a number past its end sends the frame and nothing inside it, which is how you learn you are caught up.
  • Omit since and the newest assistant message is replayed from the beginning, which is what a client that lost its place entirely needs.
  • A since naming no message of this thread replays nothing rather than restarting from the top. Restarting would repeat what you already have and hide that your cursor is stale.
No model is asked anything and nothing is written, so this costs no tokens and answers the same on a deployment with no copilot configured. What it does not do. It does not rejoin a turn that is still running, because after a disconnect there is not one: closing the request aborts the turn, the copilot stops before its next step or tool call, and what it had is stored. A reconnect therefore gives you everything you missed up to the moment the connection died. The turn does not carry on in the background.

Commands: one tool, no model

Sometimes there is nothing to work out. POST /v1/copilot/commands runs a single tool you name, as you, with no model involved and no tokens spent. It is what the console’s /screenshot and /press_key home become.
It reaches the same tool surface a chat turn drives, under the same permissions and the same approval table, so a command cannot do what the conversation would refuse. Two things differ, and both follow from your having typed the tool yourself.
  • It works with no copilot configured. A deployment with no model answers 503 on POST /v1/copilot and still runs commands. Pressing Home should not need a language model.
  • The receipt is not marked via copilot. You did this.
context and mentions work exactly as they do on the chat route, so a command typed on a phone’s page needs no deviceId. What you may drive this way: every read, plus open_app, press_key, release_device, list_runs, cancel_run and resume_run. Those are the tools whose arguments you can write out in full. tap, swipe, type_text and long_press are not, because the useful form of “tap the login button” is a sentence and the coordinates are the copilot’s job; nor is start_run, which takes a goal and a budget. Asking for one of those here is a 400 telling you to use the conversation. A tool your session has no scope for is a different answer, 403, and the message names the scope, because that remedy is an administrator rather than a rephrasing. A 200 carries {tool, isError, output: {summary, text, frameId?}} and optionally files. isError is the tool’s own verdict, not the HTTP status: a device refusing a gesture is a 200 with isError true, exactly as tool-output-error is on the stream. Send threadId and the result is appended to that thread as a tool part. That is what makes a command followed by a sentence work: the next turn sees what the command did rather than a gap where it happened. Leave it out and the command still runs, in no conversation. If a tool asks before it runs, the answer is 202 with an approval descriptor and an approvalId, and nothing has run. Show the card, then send the same tool and input again with that approvalId. The question is stored on the thread, so it can be answered minutes later and on another machine, which is why a command that asks needs a threadId. The tool and arguments are checked against what was asked: a yes to one command cannot be spent on another.

When it fails

A turn that ends badly sends error and nothing after it, with a code from a closed set. The three you will want to tell apart: model_refused, tool_failed and internal complete the set. The first three used to be one code, so a client could only ever say the vaguest of the three.

Limits

  • A turn takes at most MAX_TURN_STEPS model steps, across the requests an approval splits it into. Reaching the budget is not an error: the work done is done, and its receipts exist. A job that genuinely needs more is an agent run.
  • One message carries at most MAX_COPILOT_PROMPT_CHARS characters of text.
  • The history the model is shown is the most recent messages that fit within HISTORY_BUDGET_CHARS characters of transcript, cut at a message of yours so the transcript never opens with the copilot’s reply. Your newest message is always included. What the cut removes is not forgotten: it is folded into a short summary the copilot writes once and keeps with the thread, shown in front of the recent messages on every turn after that. A long conversation therefore still knows what it was started for and which phone it was about. The summary is regenerated only when the cut reaches past what it already covers, so it costs one extra model call when a thread first crosses the budget and nothing on the turns after. Nothing that can be rebuilt from the server goes into it: the mode, what has been remembered and which phone the chat is about are regenerated fresh every turn, so a summary of them would be a second and stale copy.
  • Screenshots travel inline up to MAX_INLINE_FILE_BYTES; above that you get the frame id.
  • The listing returns at most MAX_THREAD_ROWS threads per page, a thread read at most MAX_THREAD_MESSAGE_ROWS messages, and both refuse a larger limit rather than quietly reducing it.
  • The copilot gives up on a model that stays silent for MODEL_TIMEOUT_MS within one step, and says so as an error of its own; a model that keeps answering slowly is never cut off.
The named limits are constants in the control plane’s source, and the OpenAPI document under API carries their current values.

The older event stream

Until the console has switched, the same route also answers requests that do not ask for the UI message stream, with {prompt} in and named SSE events out (session.start, message.delta, tool.call, tool.result, approval.request, tool.denied, ping, done, error). Approvals on that wire are answered on POST /v1/copilot/approvals while the stream stays open. It goes in the release after the switch; build against the protocol above.