Skip to main content
Most of this site is written for somebody integrating against the API. This page is for the other reader: a person who signed in, wants a device, and wants to drive it without writing code first. The console is a browser face over the same control plane these docs describe. It calls the same routes with the same rules. Nothing it does is a private path, which is why every screen below can be named next to the route it calls.

Signing in

The console authenticates with an interactive session, not with an API key. The door is at console.phonebase.co, and it takes an email address and a password, or Google. Signing up confirms the address with a six digit code and creates your organization on the way in, so there is no separate step for it. Your session is the credential for every console screen. It is also the credential for the two things an API key is refused for outright: starting a purchase, and returning a device. See Two kinds of credential.

The screens that exist

A device you bought appears in the console’s device list, and in GET /v1/devices, once the order settles. Open it to reach /devices/:id. Buy, drive and return a device walks the whole loop through these screens, with the API equivalent of each step beside it.

The Control block

/devices/:id carries a Control block, and it is the part worth explaining before you use it.
  • The screen is a still frame, not a video. It is the current screen fetched from the device. There is no streaming anywhere in this product, and nothing records a session as video. See What phonebase is not.
  • Tapping the screen acts under your org’s lease. A device holds one lease at a time, across every org, and that has not changed: it is what keeps two organisations off one phone, and what the meter runs on. Inside your org the lease is shared. If an agent’s run holds it, your taps go straight to the phone under that lease and the run keeps going; the agent is told what you did on its next result and is never stopped by it. If nobody holds it, your first tap takes a lease for you, for ten minutes, and keeps it alive while you keep driving. The gesture is the browser spelling of POST /v1/devices/{deviceId}/actions, and the response says which lease it ran under and whether it was somebody else’s (coDriving).
  • The picture is the target. Click the image where you want the device tapped. The console converts where you clicked into the normalized grid described below. You never type a pixel.
  • Who’s driving is one read, GET /v1/devices/{deviceId}/driving: the run on the device, whose lease it is, and who else acted in the last minute.
  • Pause agent and Stop run are the only interruptions, and both are yours to choose: POST /v1/runs/{runId}/pause shuts the agent’s channels until you press Resume agent, and POST /v1/runs/{runId}/cancel ends the run. Acting on the screen interrupts nothing.
  • A lease you took yourself bills until it expires. Nothing watches for idleness. It expires ten minutes after your last gesture kept it alive, or you can give it back with DELETE /v1/devices/{deviceId}/lease.

A control that cannot work is disabled, and says why

Devices differ, and the console does not pretend otherwise. A control the device in front of you cannot perform is drawn disabled with its reason attached, for example Action not supported on this device. That is the same fact the API states as capabilities, and it is worth reading directly before you plan anything. One cloud phone measured today answered:
pressKey: false on a device that taps and types perfectly well is the ordinary case, not a fault. Read your own device’s block rather than carrying this one forward. Capabilities explains every field.

Coordinates are normalized, everywhere

Both axes run 0 to 1000 whatever the real screen is, with the origin at the top left in the current orientation. The device states it as coordSpace: { "width": 1000, "height": 1000 }. This is why clicking a picture works. A click at one quarter across and three quarters down a rendered image committed on a real device today at x=247, y=749. Callers never speak device pixels, and neither does the console.

What has no screen

Three surfaces are live over HTTP and have no console face at all. Do not go looking for a screen; use the API.

Two kinds of credential

An API key is refused from a browser origin. Put one in front end JavaScript and the request comes back 403 with the code api_key_from_browser_origin, whatever the key’s scopes are. The refusal is explicit rather than vague, because the audience is a developer who has just shipped a key into a browser bundle and needs to know it. Keys are for servers. Browsers get a session. If you want a browser to drive a device, the console already does that with the session it holds. Buying and returning are narrower still: both require an admin and a signed in principal, so a key cannot commit your org to a subscription and a member cannot either. See Authentication. Managing keys is the other way round. Either role does it, from a signed in session only, and the bound is ownership rather than seniority: a member mints keys and then rotates, renews and revokes the ones they minted, while an admin can act on any key in the org. None of the four asks you to reauthenticate: a re-verification step is specified and not yet built, so until it is, a signed in session is the whole of what minting a key requires.

Where to go next

Buy, drive and return a device

The whole loop, in the console and over the API, with real readings.

Authentication

The eight scopes, and how a key is narrowed to a subset of your devices.