Progress on a long action
An action call blocks until the device commits it. On the software tier that is quick. On the physical tier a mechanical actuator has to travel, press, and return, so a caller waiting on the response is waiting on a moving object. If your client sends a progress token with the request, the server reports the handoff as soon as the action is dispatched to the device:- You asked. The protocol only sends progress against a token the request carried. A request with no progress token gets no notification.
- Your transport can carry it. stdio can. The stateless HTTP path cannot, because a mid request notification needs a live bidirectional session and there is not one.
When progress is not available
Do not build a loop that depends on notifications. Build it on the recorded outcome instead, which works everywhere:- Send the action with an
idempotencyKeyyou generated. - If the response arrives, you are done. Read
statusandcommittedAt. - If it does not, call
get_receiptwith the same key.found: falsemeans the action never began. Retry with the same key.- A committed receipt means it happened. Do not retry.
- A failed receipt carrying
outcome_unknownmeans the transport to the device was lost after dispatch, so a person has to reconcile it.
When a person has to step in
Some screens are designed to stop automation, and correctly so: a sign in, a one time code, a challenge. An agent should not be trying to defeat them. The mechanism is MCP elicitation. The server asks your client’s operator to go and deal with the screen, and to confirm when they are done. The request carries the device id, a short machine stable reason, and an instruction of what to do. The person acts on the same device under the same lease; nothing is handed over and nothing is taken away.
Only an explicit confirmation counts as completed. Like progress, elicitation
needs a live bidirectional session, so it works over stdio and not over the
stateless HTTP path, where the caller sees a terminal outcome and re drives
after the person is finished.
A person’s turn inside an autonomous run
The agent tier has its own form of the same idea, and it is stronger. A worker that hits a screen needing a person reports the run asneeds_user_control. The run parks: it does not end, and while it is
parked the worker’s channel is shut on every side. A tap or a swipe is refused;
a screenshot is refused before the device is ever asked for one; and a second
run cannot be started on the same device to observe it around the first. So
what the person does while the agent waits is not served to the worker.
The only way out is a person, through resume_run or the equivalent REST
call. A worker cannot lift its own request, not by asking to resume and not by
reporting the run finished out from under the person; a worker report on a
parked run is refused until someone resumes it. (One exception is not a person:
a run still stops at its deadline, so a wait that runs long can expire. Give
the run enough wall-clock when you start it.)
A person does not have to wait to be asked. They may act on the device while
the run is driving, and the run is informed through others on its next result
rather than stopped. See Runs for how
the two hands share one phone, and for the full lifecycle.