Skip to main content
Backends differ enormously. One reaches a device by injecting input inside the operating system, another presses the glass with a mechanical actuator, and another sends hardware keyboard events with no way to see the screen at all. Rather than making you probe for that, every device declares it. Call list_devices, or get_device when you already know the id, and read the capabilities block before you plan anything.

What a device declares

The values above are one device’s answer, not a template. Read your own.

Actions

Each action field says whether the device supports something and, where it matters, by which mechanism. The mechanism changes how you should use it. pressKey is a list rather than a boolean for a reason. The tool’s parameter enumerates every key the platform has a meaningful equivalent for, and that is an upper bound across all backends. A given device supports a subset and refuses the rest.

Observations

A device with no visual observation at all, such as a hardware input dongle, declares framesAreReferenceable: false and is never forced into frame discipline. Passing observedFrame to one is refused as capability_unsupported rather than quietly ignored.

Fidelity tier

fidelityTier is the axis that decides how input reaches the device, and it takes one of three values: software, hid, and physical. See Device tiers for what each one means and what it costs you. Only the physical tier carries an extra physical block, describing abilities a mechanical actuator can have and injected input cannot: hovering above the screen without touching it, controlling press depth, and inspecting the actuator’s queued motions.
Those abilities are declared per ability rather than implied by the tier, so you can branch on each one without probing. They are declared honestly, which today means a device may report false for all of them while the device channel grows a way to carry them.
The tier is enforced when a device answers, not merely by convention: a device that is not on the physical tier carries no physical block at all, so there is nothing there to read by mistake.

Coordinates

Every tier speaks the same grid: both axes run 0 to 1000 whatever the real screen size is, with the origin at the top left, in the current orientation. coordSpace states it explicitly. Backends convert to device pixels internally. You never see pixels, except as the informational pixelSize on a screenshot, which is there to describe the image rather than to aim with. Bounds inside a ui_snapshot tree are already normalized, so a coordinate you read there can be tapped directly.

Backend kinds

backendKind names the mechanism that reaches a device. It is a closed set: The last two are declared kinds with no adapter behind them yet. A device row naming one of them is accepted, and a call against it is refused with backend_not_implemented rather than failing in some less obvious way. adb_cloud is not one of them: its adapter is live. A cloud row is reached in one of two ways, and the capability block differs between them, so read capabilities on the device rather than assuming it from the kind.
  • Over the provider’s HTTPS API. This is how the cloud phones you can buy on the spot from the /devices/add screen are reached. The frame is captured inside the guest OS (screenshot: "framebuffer"), live video comes from the provider’s own channel (stream: "webrtc", handed out by POST .../stream as transport: "provider-rtc"), and wake is supported. There is no element tree (uiTree: false), so ui_snapshot is refused with capability_unsupported; read the screen from screenshots instead. Every key in pressKey is available, and a key press is reported once the provider accepts it, without waiting for a confirmation from the device.
  • Over the device gateway tunnel. It reuses the same adb family primitives as emulator and adb_real, over a per device byte stream carried by the tunnel, which is why it declares the same software fidelity and the same capability block as they do. A cloud row reports its reachability in status, which is an answer about that device and not a statement about the adapter. Which answers you can see depends on how the row is reached. A row carried by the gateway tunnel reports offline while the tunnel is down. A row reached over the provider’s own HTTPS API reports offline when the provider named it and said it was not running, and unknown when the provider’s fleet listing failed or never mentioned it at all.
Three separate vocabularies describe a device and it is worth keeping them apart. backendKind, in this table, is the mechanism. fidelityTier is how input reaches the screen, and it is the coarser of the two: three backend kinds collapse into software. The tier names on Billing modes are a third vocabulary, for what a device costs, and that page carries the crosswalk between them and backendKind. Use backendKind to understand what you are talking to. Do not use it to infer abilities: the capability block is the authoritative answer, and a backend narrows it per device rather than widening it. A device’s backend is resolved from that device’s own record. There is no global switch, because a fleet mixes backends and a process wide setting would always be wrong for somebody’s device.

Refusal, never silence

An adapter that claims an ability it does not implement, or implements one it did not declare, fails at startup rather than at your call site. At your call site, an unsupported ability is refused with capability_unsupported. It is never accepted and quietly dropped, which would leave you believing something happened on a real device when nothing did.

What the device says it is

The capability block says what the control plane can DO to a device. The optional info block beside it says what the device IS: the four facts a person reads off Settings, About phone.
It is read once and cached, not probed to answer your call. A reading is taken the first time you acquire the device, and retried on each later acquire while it is still missing, so nothing you call pays a round trip to the device for it. A device you own but have never acquired has no reading, and will not get one until you take it. The consequence worth knowing: info can be absent on a device that will report it a minute later, and the acquire that triggers the reading can itself answer before the reading lands. resolution is the screen in device pixels. It is not the grid you aim actions at. Taps and swipes are sent in capabilities.coordSpace, a normalised 1000 by 1000 space, and converting a coordinate through the pixel rectangle instead will land it somewhere else. It is also not a screenshot’s pixelSize, which is one observation’s dimensions in whatever orientation it was taken in.

Three absences, three meanings

Missing is not one state here, and code that treats it as one will describe a phone wrongly. Which fields arrive depends on how the device is reached, not on what the hardware has. A cloud phone driven over its provider’s HTTPS API has no route that returns a shell command’s output, so nothing here comes from getprop or wm size: the model, the Android version and the operator are read from the provider’s own property API, and the screen comes from the same framebuffer measurement the tap path already uses to place a coordinate. A field is absent when that particular device did not report it, not because of the transport. The reading is taken the first time the device is acquired, and a reading that is missing a field is retaken once after a deployment, so a field can appear later without anything being asked of you. Every field is optional and the block only ever grows, so read it defensively: branch on the field you want being present, never on info having a particular shape.