ExecuServe

ExecuTorch models, served from your phone.

ExecuServe keeps compiled ExecuTorch models loaded on an Android phone and answers the OpenAI and Anthropic APIs over HTTP, in the background, for any app on the phone or, if you allow it, on your network. llama.cpp has llama-server for GGUF files; nothing equivalent existed for .pte exports, which on current phones are the fast path.

Alpha · Android 12+ on arm64 · Apache-2.0

The browser chat served by the phone: a reply from Qwen3 1.7B running on the device, with its prefill and decode rates

The browser chat, served by the phone and answering from Qwen3 1.7B running on it.

What it does

A server, not a chat app
The agents, SDKs and tools you already use call a model on your phone the way they call a cloud: OpenAI's Chat Completions, Completions and Responses, and Anthropic's Messages.
The KV cache belongs to the server
Clients resend the whole conversation, as both APIs expect; the server keeps the runtime's cache and the exact bytes of each reply, so follow-ups and tool loops still hit it.
Several models, one phone
Up to three resident, each with its own cache and endpoints, sharing one compute lane so they never halve each other's speed.
Built to stay up
A foreground service with the right locks, tested through forced deep Doze with the screen off. Where a phone maker's ROM can still stop it, the app names what to allow.
A browser chat inside
A small chat ships inside the server, so a laptop or tablet on your network can talk to the phone's models with nothing installed.
A multiplatform core
Everything but the runtime binding and the app shell is Kotlin Multiplatform and compiles for iOS on every build.

Measured on a phone

On a POCO X8 Pro Max (Dimensity 9500s), over Wi-Fi, with the minified release build.

CheckResult
Official OpenAI SDK suite16 of 16 pass
Official Anthropic SDK suite8 of 8 pass
Edge-case probe (refusals, limits, disconnects)28 of 28 pass
Screen-off soak on battery, 10 minutes20 of 20 answered
Warm second turn of a 2,000-token conversation0.47 s instead of 8.4 s
Qwen3 tool loop, second request174 of 205 prompt tokens reused
XNNPACK export against llama.cpp, Snapdragon 8 Elite1.2 to 1.6× faster decode

Method and raw results: the POCO X8 Pro Max report. The decode advantage of the XNNPACK export over llama.cpp was measured in OpenWeights, on whose ExecuTorch engine ExecuServe is built.

Quick start

An arm64 phone on Android 12 or later, adb on your computer, and an ExecuTorch export with its tokenizer.

From a clone of the repository, with the app installed:

tools/execuserve --model ~/models/Qwen3-1.7B-8da4w-gptq-2k.pte
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1
export OPENAI_API_KEY=<the key it prints>

Then any OpenAI client:

from openai import OpenAI

client = OpenAI()
reply = client.chat.completions.create(
    model="qwen3-1.7b",
    messages=[{"role": "user", "content": "Hello"}],
)

The script pushes the model, starts the server, forwards the port over adb, and prints the base URL and key. Anthropic SDKs, Open WebUI, the codex CLI and other apps on the same phone work too; the README has the settings for each.

The API

OpenAI's and Anthropic's shapes, error codes and streaming, checked with their official SDKs.

POST /v1/chat/completionsStreaming, tools and tool calls, reasoning, usage with cached tokens
POST /v1/responsesItems, function calls, typed stream events, previous_response_id
POST /v1/messagesAnthropic's Messages API: tool_use, thinking blocks, its stream events
POST /v1/completionsA raw prompt, no template
GET /v1/modelsInstalled models with context length, residency and capabilities
GET /metricsPrometheus counters, gauges and quantiles
/models/{id}/v1The model, inference and status routes scoped to one model, for hosting several at once

Private by construction

No account, no analytics, no crash reporter and no server of ours. Prompts and replies are answered on the phone and never sent to us; the run history keeps figures, never what was asked. Every request needs a key, loopback included, and the browser chat is served under a strict Content-Security-Policy. The privacy policy says exactly what stays and what leaves.

Open source, and what comes next

Apache-2.0, on GitHub. The server core is Kotlin Multiplatform and already compiles for iOS; next are the iOS app, the other ExecuTorch backends (Vulkan, QNN, MediaTek), server-side tools, vision input, and TLS for network mode. Not affiliated with, endorsed by or sponsored by the PyTorch Foundation or Meta.