> ## Documentation Index
> Fetch the complete documentation index at: https://docs.phoneuse.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Progress and takeover

> Long running actions, and the moments a person has to step in.

Two facts about driving real hardware shape this page. An action takes real
time, because something physical has to happen. And some screens cannot be
solved by an agent at all, because they are asking for a person.

Phonebase handles both, and how well it handles them depends on your
transport.

## Progress on a long action

An action call blocks until the **device** commits it. On the software tier
that is quick. On the physical tier a mechanical actuator has to travel, press,
and return, so a caller waiting on the response is waiting on a moving object.

If your client sends a progress token with the request, the server reports the
handoff as soon as the action is dispatched to the device:

```text theme={null}
dispatched to black-arm-01; awaiting device commit
```

That notification says the request reached the device and physical execution
is in flight. It does not say the action succeeded. The success answer is the
result itself, and it arrives when the device confirms.

Two conditions have to hold for any of this to reach you.

1. **You asked.** The protocol only sends progress against a token the request
   carried. A request with no progress token gets no notification.
2. **Your transport can carry it.** stdio can. The stateless HTTP path cannot,
   because a mid request notification needs a live bidirectional session and
   there is not one.

A transport that cannot deliver the notification does not fail your action
over it. Progress is a transport capability, never a promise the tool surface
makes.

## When progress is not available

Do not build a loop that depends on notifications. Build it on the recorded
outcome instead, which works everywhere:

1. Send the action with an `idempotencyKey` you generated.
2. If the response arrives, you are done. Read `status` and `committedAt`.
3. If it does not, call `get_receipt` with the same key.
   * `found: false` means the action never began. Retry with the **same** key.
   * A committed receipt means it happened. Do not retry.
   * A failed receipt carrying `outcome_unknown` means the transport to the
     device was lost after dispatch, so a person has to reconcile it.

This path survives a lost response, a client restart, and a transport with no
notification channel at all. See [Reading a trace](/getting-started/trace).

## When a person has to step in

Some screens are designed to stop automation, and correctly so: a sign in, a
one time code, a challenge. An agent should not be trying to defeat them.

The mechanism is MCP elicitation. The server asks **your client's** operator to
go and deal with the screen, and to confirm when they are done. The request
carries the device id, a short machine stable reason, and an instruction of
what to do. The person acts on the same device under the same lease; nothing is
handed over and nothing is taken away.

```json theme={null}
{
  "message": "[black-arm-01] Complete the sign in on the device (reason: login_required)",
  "requestedSchema": {
    "type": "object",
    "properties": { "done": { "type": "boolean", "title": "Done" } },
    "required": ["done"]
  }
}
```

<Warning>
  No secret ever travels this channel. The person acts **on the device**, not
  in a chat window. Phonebase asks a human to go and type a password into the
  phone in front of them; it never asks anyone to send it one.
</Warning>

The wait is generous, because a person unlocking a device and solving a
challenge takes minutes rather than the seconds a normal remote call is given.
It is still bounded, because an elicitation that never resolves would hold a
lease open behind it.

Every non confirmation fails closed:

| What happened | Outcome |
| - | - |
| The operator confirmed they finished | `completed` |
| The client cannot do elicitation at all | `unavailable` |
| Declined, cancelled, timed out, or the transport failed | `declined` |

Only an explicit confirmation counts as completed. Like progress, elicitation
needs a live bidirectional session, so it works over stdio and not over the
stateless HTTP path, where the caller sees a terminal outcome and re drives
after the person is finished.

## A person's turn inside an autonomous run

The agent tier has its own form of the same idea, and it is stronger.

A worker that hits a screen needing a person reports the run as
`needs_user_control`. The run **parks**: it does not end, and while it is
parked the worker's channel is shut on every side. A tap or a swipe is refused;
a screenshot is refused before the device is ever asked for one; and a second
run cannot be started on the same device to observe it around the first. So
what the person does while the agent waits is not served to the worker.

The only way out is a person, through `resume_run` or the equivalent REST
call. A worker cannot lift its own request, not by asking to resume and not by
reporting the run finished out from under the person; a worker report on a
parked run is refused until someone resumes it. (One exception is not a person:
a run still stops at its deadline, so a wait that runs long can expire. Give
the run enough wall-clock when you start it.)

A person does not have to wait to be asked. They may act on the device while
the run is driving, and the run is informed through `others` on its next result
rather than stopped. See [Runs](/agent/runs#driving-alongside-a-run) for how
the two hands share one phone, and for the full lifecycle.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.