An agent driving a real device has to answer one question after every call:
did that actually happen? Phonebase answers it with a recorded outcome rather
than with an acknowledgement, and gives you a way to look the answer up again
after the fact.
Three things carry that guarantee: the typed result, the idempotency key, and
the observed frame.
The typed result
Every successful call returns typed structuredContent, shaped by the output
schema that tool publishes. The same value is also written into a text block
for older consumers, so a client that reads either one sees the same data.
For an action, that value is an outcome:
Read it precisely.
status: "committed" means the device applied the action. It is not an
acknowledgement that the request arrived.
committedAt is when the backend confirmed the commit, on the control
plane’s clock.
deduplicated: true means this response was replayed from the ledger and
the gesture did not happen again.
Failures never appear in this shape. A failed action is an error result, and
errors have their own envelope.
The idempotency key
idempotencyKey is a short printable ASCII request id that you generate and
pass with an action. Its exact length and character bounds are published in
the tool reference. It is optional, and it is worth passing
every time.
With a key of your own, a retry of a request whose response you lost replays
what was recorded instead of pressing the glass a second time. Without one the
control plane generates a key so the action is still recorded, but it cannot
recognise your retry as a retry, because it never saw your first attempt’s
key.
Two limits are worth knowing before you build a retry loop:
- A key that failed terminally will not replay into a fresh attempt. Reusing
it replays the recorded failure, and the error tells you to retry under a new
key. Idempotency protects you from double execution; it is not a way to make
a failed action run again.
- Some outcomes are genuinely unknown. The error set includes a case for an
action whose result the control plane could not establish. Treat it as
neither done nor not done, and resolve it by reading the receipt rather than
by retrying blind.
Keys beginning with agent: are refused. That prefix belongs to the control
plane’s own runs, and a caller able to write there could pre create receipts
for a run’s future actions and have every one of them replay instead of
execute.
Reuse a key only for the same intended action. A different tap under the
same key is a conflict, not a retry.
Looking an outcome up later
When a response is lost, do not retry blind. Ask what happened:
found: false is a normal answer, not an error. It means no receipt exists
under that key, so the action never began, and the same key is safe to
retry.
Two fields deserve attention when a receipt is not simply committed:
reason: "interrupted_by_restart" means the control plane restarted while
the action was in flight, so the request was closed as failed rather than
guessed either way.
reason: "outcome_unknown" means the transport to the device was lost after
dispatch. Whether the device acted is genuinely unknown.
Both arrive with requiresManualReview: true. That flag exists because
guessing is worse than admitting the gap: a person has to reconcile what the
device actually did.
get_receipt reads receipts, so it is available to a read only credential. A
lease id and fencing token are never projected into it, because they are
bearer credentials rather than audit fields.
The observed frame
observedFrame is the other half of the trace, and it works forwards rather
than backwards.
A screenshot issues a frameId. Pass that id back on the next action and the
action is refused unless the frame is still the device’s latest. A refusal comes
back as stale_frame, telling you the screen moved between deciding and acting.
There is no clock on it: a frame does not age out, so however long you spend
deciding does not matter.
This is what answers the classic worry of a vision driven loop: deciding on a
picture and then tapping a screen somebody else has changed. It answers it by
telling you: every action comes back with others, which says
how many actions another principal took since the frame you named, when, and
whether by a person or an agent. An agent that names a frame is also refused
when that count is above zero, by default; pass refuseIfOthersActed: false to
be told without being stopped.
Naming a frame that was never issued is still an error, and so is naming one so
far back that the control plane no longer holds it.
Not every device supports it. A device whose capabilities report
framesAreReferenceable: false has no visual observation to bind to, and
passing observedFrame there is refused as capability_unsupported.
How to consume an error
Errors arrive with isError: true and come in exactly two shapes, so there is
no ambiguity about how to parse one.
- A device or control plane error. The text block is a JSON envelope with
a stable
code, a message, and retryable, plus any fields specific to
that error. Branch on code.
- A schema rejection. The text block is prose, and it will not parse as
JSON. Your parameters were wrong. Fix them before resending.
So: try to parse the text block. If it parses, branch on the code. If it does
not, you have an argument problem.
Errors never carry structuredContent. That is deliberate, because a standard
MCP client validates any structuredContent it sees against the tool’s output
schema with no exemption for errors, and an envelope there would break it.
The full list of codes is in the error reference, and the
common ones are explained in Troubleshooting.
Why usage and evidence agree
The receipt ledger is the system of record. Usage is derived from it rather
than counted separately, so what you are billed for and what you can look up
with get_receipt are the same events. See Metering units.