mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Dev Tools

Mobile-MCP: How Model Context Protocol Servers Turn iOS and Android Devices Into Agent Tool Endpoints

MCP server architecture for mobile automation: accessibility snapshots, coordinate-based taps, and device abstraction for native iOS/Android apps.

Source: github.com
Mobile-MCP: How Model Context Protocol Servers Turn iOS and Android Devices Into Agent Tool Endpoints

Mobile-MCP is a Model Context Protocol server that exposes iOS and Android devices as tool endpoints for LLM agents. Instead of writing platform-specific automation scripts, agents call standardized MCP tools to query accessibility trees, tap coordinates, and navigate native apps without knowing UIKit from Jetpack Compose.

The project (7,846 stars, trending #9 in TypeScript) works with Claude Code, Codex, Gemini, GitHub Copilot, and any MCP-compatible client. It supports emulators, simulators, and real devices through a single interface. Mobile Next Cloud offers production-ready remote device infrastructure with the same tool contract.

Why This Matters Now

MCP adoption is accelerating across agent frameworks, but most implementations target web browsers or local filesystems. Mobile apps remain locked behind platform-specific APIs: XCTest for iOS, UIAutomator for Android, and fragile coordinate-based scripts that break with every UI update.

Mobile-MCP solves this by abstracting device interaction into a protocol layer. Agents see a unified schema for accessibility snapshots and input events. The server handles platform translation, session isolation, and state cleanup. This turns mobile devices into first-class tool endpoints in multi-step agent workflows.

Architecture: Device Abstraction Through MCP

Mobile-MCP sits between the agent runtime and device automation drivers. The server exposes MCP tools for:

  • Accessibility snapshots: Structured JSON representing the current view hierarchy (labels, IDs, bounds, states).
  • Coordinate taps: Pixel-based input when accessibility data is insufficient or unavailable.
  • Navigation actions: Swipe, scroll, type, and app lifecycle commands.
  • Screenshot capture: Visual context for vision-enabled models.

The server translates MCP tool calls into platform-specific commands:

  • iOS: Uses xcrun simctl for simulators or WebDriverAgent for real devices, querying accessibility trees via XCTest APIs.
  • Android: Uses ADB for emulators or real devices, querying view hierarchies via UIAutomator dump.

Agents never see the platform boundary. They call get_accessibility_tree and receive a normalized JSON structure regardless of whether the target is an iPhone 15 simulator or a Pixel 8 emulator.

Protocol Boundary: Accessibility vs. Coordinates

Mobile-MCP offers two interaction modes, and the choice affects reliability and maintenance burden.

Accessibility-based actions use semantic element IDs or labels. The agent calls tap_element with an accessibility identifier, and the server resolves it to coordinates at runtime. This survives layout changes but requires apps to expose accessibility metadata.

Coordinate-based taps use pixel positions derived from screenshots. The agent analyzes the image, identifies the target UI element, and calls tap_coordinate with x/y values. This works for apps with poor accessibility labeling but breaks when screen dimensions or layouts change.

The server chooses the mode based on what the agent requests. Most production workflows start with accessibility queries and fall back to coordinate taps only when semantic data is missing.

ModeReliabilityMaintenanceUse Case
Accessibility-basedHigh (survives layout changes)Low (uses stable IDs)Apps with proper a11y labels
Coordinate-basedMedium (breaks on layout shift)High (requires vision model)Legacy apps, games, custom UI
HybridHigh (best of both)Medium (fallback logic)Production agent workflows

Device Session Management in Multi-Tenant Environments

Mobile Next Cloud runs Mobile-MCP against real iOS and Android hardware in a shared infrastructure. This introduces session isolation and state cleanup challenges.

Session isolation: Each agent workflow gets an exclusive device lock. The server spawns a new device session, installs required apps, and tears down state after the workflow completes. Concurrent agents cannot access the same device, but the scheduler can allocate different devices from the pool.

State cleanup: After each session, the server:

  • Uninstalls test apps to prevent data leakage.
  • Clears app caches and local storage.
  • Resets accessibility settings to defaults.
  • Captures logs and screenshots for observability.

Failure modes: If an agent crashes mid-workflow, the server detects the orphaned session via timeout and forces a device reset. This prevents one bad workflow from poisoning the device pool.

Local setups skip most of this complexity. The server assumes single-user access and leaves state management to the developer.

Tool Call Flow: From Agent Request to Device Action

Here’s how an agent interacts with a mobile app through Mobile-MCP:

  1. Agent calls get_accessibility_tree: The MCP client sends a tool request to the server.
  2. Server queries the device: For iOS, runs xcrun simctl io booted accessibility dump. For Android, runs adb shell uiautomator dump.
  3. Server normalizes the response: Converts platform-specific XML or plist output into a JSON schema the agent understands.
  4. Agent analyzes the tree: Identifies the target element (e.g., a “Submit” button with ID submit_btn).
  5. Agent calls tap_element: Sends the element ID back to the server.
  6. Server resolves coordinates: Looks up the element’s bounding box in the cached accessibility tree.
  7. Server executes the tap: For iOS, uses WebDriverAgent to send a touch event. For Android, uses adb shell input tap.
  8. Agent requests a new snapshot: Calls get_accessibility_tree again to verify the UI state changed.

The entire flow is synchronous from the agent’s perspective. The server handles retries, coordinate translation, and platform quirks.

Code Snippet: Querying and Tapping an Element

This example shows how an agent might use Mobile-MCP tools to log into a mobile app. The agent runtime (e.g., Claude Code) calls these tools via the MCP protocol.

// Agent workflow pseudocode (not actual Mobile-MCP server code)

// Step 1: Get the current screen state
const tree = await mcp.callTool("get_accessibility_tree", {
  deviceId: "iphone-15-sim"
});

// Step 2: Find the username field
const usernameField = tree.elements.find(
  el => el.label === "Username" && el.type === "textField"
);

// Step 3: Tap and type
await mcp.callTool("tap_element", {
  deviceId: "iphone-15-sim",
  elementId: usernameField.id
});

await mcp.callTool("type_text", {
  deviceId: "iphone-15-sim",
  text: "user@example.com"
});

// Step 4: Submit the form
const submitButton = tree.elements.find(
  el => el.label === "Submit"
);

await mcp.callTool("tap_element", {
  deviceId: "iphone-15-sim",
  elementId: submitButton.id
});

// Step 5: Verify success
const newTree = await mcp.callTool("get_accessibility_tree", {
  deviceId: "iphone-15-sim"
});

const successMessage = newTree.elements.find(
  el => el.label.includes("Welcome")
);

The agent never writes iOS or Android code. It just calls MCP tools and inspects the returned JSON.

Observability and Debugging

Mobile-MCP logs every tool call, device command, and response payload. This is critical for debugging flaky workflows.

Structured logs include:

  • Tool name and parameters
  • Device ID and platform
  • Command executed (e.g., adb shell input tap 500 800)
  • Response time and status
  • Screenshot hash for visual correlation

Failure diagnostics: When a tap fails, the server captures:

  • The accessibility tree before and after the action
  • A screenshot showing the actual UI state
  • The element bounds and calculated tap coordinates
  • Any platform errors (e.g., “Element not found” from UIAutomator)

This data feeds into retry logic or surfaces to the agent for re-planning.

Deployment Shapes

Local development: Install Mobile-MCP via npm, point it at local simulators or emulators, and connect your MCP client. The server runs as a long-lived process on your machine.

CI/CD pipelines: Spin up Mobile-MCP in a Docker container alongside Android emulators or iOS simulators (requires macOS runners for iOS). Agents run end-to-end tests against ephemeral devices.

Mobile Next Cloud: Use the hosted service for real device access. The server runs in their infrastructure, and you connect via API keys. No local setup, but you pay per device-minute.

Self-hosted cloud: Deploy Mobile-MCP on your own infrastructure with a device farm (e.g., AWS Device Farm, BrowserStack). Requires custom session management and device provisioning logic.

Security Boundaries

Mobile-MCP has full control over the target device. An agent can install apps, read files, and execute arbitrary shell commands (depending on server configuration).

Mitigation strategies:

  • Run the server in a sandboxed environment (container, VM, or dedicated device pool).
  • Restrict tool permissions via MCP server config (e.g., disable install_app or execute_shell).
  • Use device-level restrictions (e.g., iOS Guided Access, Android kiosk mode).
  • Audit all tool calls and flag suspicious patterns (e.g., repeated failed login attempts).

Mobile Next Cloud enforces tenant isolation at the device level. Each session runs on a dedicated device, and state is wiped between sessions. But the server itself is a privileged component, so treat it like a production database.

Likely Failure Modes

Accessibility tree drift: Apps update their UI, and element IDs change. Agents retry with stale identifiers and fail. Solution: Use fuzzy matching on labels or fall back to coordinate taps.

Coordinate misalignment: Screenshots are captured at one resolution, but the device renders at another (e.g., retina displays). Taps land in the wrong spot. Solution: Normalize coordinates to logical pixels.

Session timeouts: Long-running workflows exceed the device lock timeout. The server kills the session mid-task. Solution: Implement heartbeat pings or increase timeout limits.

Platform API changes: iOS or Android updates break the underlying automation drivers (XCTest, UIAutomator). Solution: Pin driver versions and test against beta OS releases.

Concurrency limits: More agents request devices than the pool can handle. Workflows queue or fail. Solution: Scale the device pool or implement priority-based scheduling.

Technical Verdict

Use Mobile-MCP when:

  • You need agents to interact with native mobile apps without writing platform-specific code.
  • You want a single tool contract for iOS, Android, simulators, emulators, and real devices.
  • You’re building multi-step workflows that combine mobile automation with other MCP tools (browsers, APIs, databases).
  • You need production-ready device infrastructure and don’t want to manage a device farm.

Avoid Mobile-MCP when:

  • You only automate web apps (use browser-based MCP servers instead).
  • You need sub-second latency for real-time interactions (accessibility queries add overhead).
  • Your apps lack accessibility metadata and you can’t rely on coordinate-based fallbacks.
  • You require platform-specific features that the abstraction layer doesn’t expose (e.g., iOS Shortcuts, Android Intents).

Mobile-MCP turns mobile devices into tool endpoints for agents. The abstraction works because it exposes the right primitives: accessibility trees for semantic actions, coordinates for visual fallbacks, and session isolation for multi-tenant safety. The plumbing is straightforward, but the operational complexity lives in device provisioning, state cleanup, and failure recovery.