POST /v1/copilot, and four routes for
the threads it writes. All five need a signed in person: a copilot acts as
somebody, and an API key has nobody behind it.
Threads live on the server
A conversation is a thread, owned by the person who opened it, inside their organisation. The server keeps it: every message you send, every message the copilot writes, every tool it called and what came back, and the screenshots it took along the way. Three things follow from that.- The copilot remembers. Each turn is shown the thread so far, within a budget (see Limits), so you can say “now open the second one” without repeating yourself.
- Your client’s history is not trusted. A request carries the client’s copy of the transcript, and only its last message is read. What the copilot is shown is the stored thread. Two browsers open on one thread see the same past, and nobody can put words in the copilot’s mouth by editing a payload.
- A thread is private. Another person in your organisation cannot list it, open it, rename it or delete it. To them it answers 404, the same as a thread that never existed. There is no sharing yet.
A new thread takes its title from the first message you send, cut to the
length
THREAD_TITLE_MAX_CHARS allows.
Reopening a thread
GET /v1/copilot/threads/{id} answers with the thread and its newest messages,
oldest first, so the page ends where the conversation stands. It is paged
because a thread is heavy: every tool result carries its full text and every
screenshot its bytes.
The response also carries mode: the mode saved on this thread, or null if
none was chosen. A second browser can show the same choice when it reopens the
thread.
limitis how many messages,DEFAULT_THREAD_MESSAGE_ROWSwhen omitted and at mostMAX_THREAD_MESSAGE_ROWS.nextBeforein the response is the oldest message on the page, when older ones exist. Send it back asbeforeto get the page before it.- Inline screenshots are carried within
MAX_INLINE_PAGE_BYTESof decoded image data per response (about a third more on the wire, as base64 inside JSON). The fit is greedy, newest first: each screenshot is kept if it fits what is left, so a large one that does not fit is left out while a smaller older one may still be carried.files=noneleaves every one out. Either wayomittedFilessays how many were left out, and theframeIdon the tool output beside each one still names the picture for the frames API.
messages to your chat state as they are. A tool part in state
approval-requested on the last message is a question still waiting.
The protocol
POST /v1/copilot answers in the UI message stream when your request lists
application/vnd.phonebase.ui-message-stream in Accept. The body is the one
the AI SDK’s default chat transport sends:
id is the thread. A new id opens a new thread owned by you. messages is
your transcript; only the last message is read, and only its text parts.
User file attachments are rejected with invalid_argument. The copilot still
returns screenshots it observes on the phone.
Which phone the turn is about
A person on the page of one phone who types “open Settings” has already named the phone. Send what you know beside the message and the copilot will not ask them to say it again:context and mentions are optional, and both go on either body shape.
mentions carries device ids, not the names in the sentence, so renaming a
phone can never re-point a message somebody already sent.
The copilot picks the phone in this order and stops at the first hit:
context.pinned.deviceId, a phone the person pinned to this chat.mentions, when it names exactly one.context.device.id, the page they are on.- The device this thread last acted on, read from its own tool calls.
- The only device you have, if you have one.
- Otherwise a
data-device-choicecard.
- The copilot is told, and it acts. A tool call that needs a device and names none is given the one the context decided, so a sentence with no phone in it still reaches the right phone, and the call you see carries the id.
contextis never replayed. It describes where somebody is now, so only this request’s copy is read. The stored thread keeps none of it. What survives between turns is which device the thread actually touched, which is a fact rather than a claim about a page.- An id you cannot address is a 400, naming the field. Every field that is
wrong is named in the one answer, so you never fix them one round trip at a
time.
context.deviceis{"id": ...}, an object; a bare string is refused rather than ignored.
context.device and the rest may carry more than the id (114 draws them with
a name, a tier and an online flag), and extra keys are accepted. Only the ids
are read. The name, tier and online state the copilot works from are this
service’s own, so a page that has gone stale cannot tell it a phone is up.
The response is text/event-stream, with the header
x-vercel-ai-ui-message-stream: v1 so a transport can tell it apart from the
older event stream without reading the body. Each frame is data: followed
by one JSON chunk, and the stream ends with data: [DONE].
The chunks, in the order a turn produces them:
Between chunks the stream may carry SSE comment lines as keep-alives. An
event-source parser ignores them.
Closing the request stops the turn. The copilot finishes the tool call it is
on, if one is running, and stops before its next step or tool call. What it
wrote so far is kept, and the stored message ends with
finishReason: "stop".
Approvals
Two things ask you first: installing an app, and starting an agent run that goes on acting by itself after the turn ends. Everything else, looking, taking and letting go of a device, and every gesture, simply happens because you asked for it. A card in front of every tap is not consent; it is a clickthrough, and it teaches people to approve without reading. The question rides the stream and ends the turn:- The copilot sends
tool-input-availablefor the call, thentool-approval-requestwith anapprovalIdand a descriptor: the tool, its arguments, one sentence to show, andexpiresAt. The sentence is built from the tool and its arguments, never from anything the model wrote, so a card cannot describe a tap as a screenshot. - The stream ends with
finish{finishReason: "tool-calls"}. The stored assistant message holds the pending question, so the answer may arrive on any machine, minutes later. - You answer by sending the same thread again, with that assistant message
last and its tool part in state
approval-responded, carryingapproval: {id, approved, reason?, input?, descriptor?}. This is what the AI SDK’saddToolApprovalResponseproduces, plus the two fields below. - The next stream continues the same assistant message: its
startnames the samemessageId. On a yes the copilot looks at the screen first, then runs the call, then the calls the model had queued behind it, then decides again. On a no it sendstool-output-deniedand the turn ends: a copilot that carried on after a refusal, working around it, would make its approvals meaningless.
What the card says
The descriptor decides which card to draw and why, so the client never has to work it out from the tool name:
A descriptor written before these fields existed carries neither
kind nor
reason. Read a missing kind as light.
One card, several phones
Ask for the same thing on three phones and you get one question, not three. The descriptor’scalls lists every call it covers, one yes runs all of them,
and each still runs as its own call and leaves its own receipt. A dangerous
operation is never folded into that list: it always gets its own card, so a
yes to three of something can never have been a yes to something else standing
beside them.
Editing before you agree
On a plan card somebody may change the goal, the step ceiling or the time limit before saying yes. Send the changed fields asapproval.input.
approval.input after the call has run, so a thread
read back later still shows the edit the call ran with.
If you echo approval.descriptor.reason back, it has to be the code the
question was asked with, or the answer is refused: consent shown for one
reason is not consent for another. Echoing nothing is accepted.
A question that nobody answers within APPROVAL_TIMEOUT_MS expires: the answer
is refused with tool-output-error on the call. Sending a new message while a
question is pending closes the question the same way.
Only the person the question was asked of can answer it, because the thread is
theirs. An answer to a question that is not pending, or that this thread never
asked, is 404.
How much it does on its own
Every chat runs in one of three modes. The product default isauto, and that
is a deliberate choice rather than a convenience: a card in front of every
action is not consent, it is a clickthrough.
Send
mode on a chat request to change that chat’s mode. It takes effect on
the turn that sets it and is saved with the thread. An approval authorizes
only the pending action. There are no remembered grants or organisation-wide
automation defaults.
What is never turned off
No mode bypasses any of these.- A tool that declares
requiresUserInteractionin its own MCP metadata. This is enforced at every call. Historicalinstall_appapproval cards remain readable, but approving one cannot restore retired installation. - A text message to somebody this chat has not written to before.
- Doing something to several phones in one go.
- Creating a schedule.
- Money, API keys and team membership, which are never done on your behalf at all. Those are not cards; they are refusals with a link.
Picking a cut stream back up
A phone that went to the background, a tab switched away from, a network that dropped. Reconnect toGET /v1/copilot/threads/{id}/stream and the rest of
the message arrives as the same chunks, under the same header, ending with the
same data: [DONE].
sinceis themessageIdfrom thestartchunk you were reading.afterPartis how many of its parts you already hold.0, the default, replays the whole message; a number past its end sends the frame and nothing inside it, which is how you learn you are caught up.- Omit
sinceand the newest assistant message is replayed from the beginning, which is what a client that lost its place entirely needs. - A
sincenaming no message of this thread replays nothing rather than restarting from the top. Restarting would repeat what you already have and hide that your cursor is stale.
Commands: one tool, no model
Sometimes there is nothing to work out.POST /v1/copilot/commands runs a
single tool you name, as you, with no model involved and no tokens spent. It is
what the console’s /screenshot and /press_key home become.
- It works with no copilot configured. A deployment with no model answers
503 on
POST /v1/copilotand still runs commands. Pressing Home should not need a language model. - The receipt is not marked
via copilot. You did this.
context and mentions work exactly as they do on the chat route, so a
command typed on a phone’s page needs no deviceId.
What you may drive this way: every read, plus open_app, press_key,
release_device, list_runs, cancel_run and resume_run. Those are the
tools whose arguments you can write out in full. tap, swipe, type_text
and long_press are not, because the useful form of “tap the login button” is
a sentence and the coordinates are the copilot’s job; nor is start_run, which
takes a goal and a budget. Asking for one of those here is a 400 telling you to
use the conversation. A tool your session has no scope for is a different
answer, 403, and the message names the scope, because that remedy is an
administrator rather than a rephrasing.
A 200 carries {tool, isError, output: {summary, text, frameId?}} and
optionally files. isError is the tool’s own verdict, not the HTTP status: a
device refusing a gesture is a 200 with isError true, exactly as
tool-output-error is on the stream.
Send threadId and the result is appended to that thread as a tool part. That
is what makes a command followed by a sentence work: the next turn sees what
the command did rather than a gap where it happened. Leave it out and the
command still runs, in no conversation.
If a tool asks before it runs, the answer is 202 with an approval descriptor
and an approvalId, and nothing has run. Show the card, then send the same
tool and input again with that approvalId. The question is stored on the
thread, so it can be answered minutes later and on another machine, which is
why a command that asks needs a threadId. The tool and arguments are checked
against what was asked: a yes to one command cannot be spent on another.
When it fails
A turn that ends badly sendserror and nothing after it, with a code from a
closed set. The three you will want to tell apart:
model_refused, tool_failed and internal complete the set. The first three
used to be one code, so a client could only ever say the vaguest of the three.
Limits
- A turn takes at most
MAX_TURN_STEPSmodel steps, across the requests an approval splits it into. Reaching the budget is not an error: the work done is done, and its receipts exist. A job that genuinely needs more is an agent run. - One message carries at most
MAX_COPILOT_PROMPT_CHARScharacters of text. - The history the model is shown is the most recent messages that fit within
HISTORY_BUDGET_CHARScharacters of transcript, cut at a message of yours so the transcript never opens with the copilot’s reply. Your newest message is always included. What the cut removes is not forgotten: it is folded into a short summary the copilot writes once and keeps with the thread, shown in front of the recent messages on every turn after that. A long conversation therefore still knows what it was started for and which phone it was about. The summary is regenerated only when the cut reaches past what it already covers, so it costs one extra model call when a thread first crosses the budget and nothing on the turns after. Nothing that can be rebuilt from the server goes into it: the mode, what has been remembered and which phone the chat is about are regenerated fresh every turn, so a summary of them would be a second and stale copy. - Screenshots travel inline up to
MAX_INLINE_FILE_BYTES; above that you get the frame id. - The listing returns at most
MAX_THREAD_ROWSthreads per page, a thread read at mostMAX_THREAD_MESSAGE_ROWSmessages, and both refuse a largerlimitrather than quietly reducing it. - The copilot gives up on a model that stays silent for
MODEL_TIMEOUT_MSwithin one step, and says so as an error of its own; a model that keeps answering slowly is never cut off.
The older event stream
Until the console has switched, the same route also answers requests that do not ask for the UI message stream, with{prompt} in and named SSE events out
(session.start, message.delta, tool.call, tool.result,
approval.request, tool.denied, ping, done, error). Approvals on that
wire are answered on POST /v1/copilot/approvals while the stream stays open.
It goes in the release after the switch; build against the protocol above.