Skip to main content
An agent driving a real device has to answer one question after every call: did that actually happen? Phonebase answers it with a recorded outcome rather than with an acknowledgement, and gives you a way to look the answer up again after the fact. Three things carry that guarantee: the typed result, the idempotency key, and the observed frame.

The typed result

Every successful call returns typed structuredContent, shaped by the output schema that tool publishes. The same value is also written into a text block for older consumers, so a client that reads either one sees the same data. For an action, that value is an outcome:
Read it precisely.
  • status: "committed" means the device applied the action. It is not an acknowledgement that the request arrived.
  • committedAt is when the backend confirmed the commit, on the control plane’s clock.
  • deduplicated: true means this response was replayed from the ledger and the gesture did not happen again.
Failures never appear in this shape. A failed action is an error result, and errors have their own envelope.

The idempotency key

idempotencyKey is a short printable ASCII request id that you generate and pass with an action. Its exact length and character bounds are published in the tool reference. It is optional, and it is worth passing every time. With a key of your own, a retry of a request whose response you lost replays what was recorded instead of pressing the glass a second time. Without one the control plane generates a key so the action is still recorded, but it cannot recognise your retry as a retry, because it never saw your first attempt’s key. Two limits are worth knowing before you build a retry loop:
  • A key that failed terminally will not replay into a fresh attempt. Reusing it replays the recorded failure, and the error tells you to retry under a new key. Idempotency protects you from double execution; it is not a way to make a failed action run again.
  • Some outcomes are genuinely unknown. The error set includes a case for an action whose result the control plane could not establish. Treat it as neither done nor not done, and resolve it by reading the receipt rather than by retrying blind.
Keys beginning with agent: are refused. That prefix belongs to the control plane’s own runs, and a caller able to write there could pre create receipts for a run’s future actions and have every one of them replay instead of execute.
Reuse a key only for the same intended action. A different tap under the same key is a conflict, not a retry.

Looking an outcome up later

When a response is lost, do not retry blind. Ask what happened:
found: false is a normal answer, not an error. It means no receipt exists under that key, so the action never began, and the same key is safe to retry. Two fields deserve attention when a receipt is not simply committed:
  • reason: "interrupted_by_restart" means the control plane restarted while the action was in flight, so the request was closed as failed rather than guessed either way.
  • reason: "outcome_unknown" means the transport to the device was lost after dispatch. Whether the device acted is genuinely unknown.
Both arrive with requiresManualReview: true. That flag exists because guessing is worse than admitting the gap: a person has to reconcile what the device actually did. get_receipt reads receipts, so it is available to a read only credential. A lease id and fencing token are never projected into it, because they are bearer credentials rather than audit fields.

The observed frame

observedFrame is the other half of the trace, and it works forwards rather than backwards. A screenshot issues a frameId. Pass that id back on the next action and the action is refused unless the frame is still the device’s latest. A refusal comes back as stale_frame, telling you the screen moved between deciding and acting. There is no clock on it: a frame does not age out, so however long you spend deciding does not matter. This is what answers the classic worry of a vision driven loop: deciding on a picture and then tapping a screen somebody else has changed. It answers it by telling you: every action comes back with others, which says how many actions another principal took since the frame you named, when, and whether by a person or an agent. An agent that names a frame is also refused when that count is above zero, by default; pass refuseIfOthersActed: false to be told without being stopped. Naming a frame that was never issued is still an error, and so is naming one so far back that the control plane no longer holds it. Not every device supports it. A device whose capabilities report framesAreReferenceable: false has no visual observation to bind to, and passing observedFrame there is refused as capability_unsupported.

How to consume an error

Errors arrive with isError: true and come in exactly two shapes, so there is no ambiguity about how to parse one.
  1. A device or control plane error. The text block is a JSON envelope with a stable code, a message, and retryable, plus any fields specific to that error. Branch on code.
  2. A schema rejection. The text block is prose, and it will not parse as JSON. Your parameters were wrong. Fix them before resending.
So: try to parse the text block. If it parses, branch on the code. If it does not, you have an argument problem. Errors never carry structuredContent. That is deliberate, because a standard MCP client validates any structuredContent it sees against the tool’s output schema with no exemption for errors, and an envelope there would break it. The full list of codes is in the error reference, and the common ones are explained in Troubleshooting.

Why usage and evidence agree

The receipt ledger is the system of record. Usage is derived from it rather than counted separately, so what you are billed for and what you can look up with get_receipt are the same events. See Metering units.