ExecuTorch models, served from your phone.
ExecuServe keeps compiled ExecuTorch models loaded on an Android phone and answers the OpenAI and Anthropic APIs over HTTP, in the background, for any app on the phone or, if you allow it, on your network. llama.cpp has llama-server for GGUF files; nothing equivalent existed for .pte exports, which on current phones are the fast path.
Alpha · Android 12+ on arm64 · Apache-2.0

The browser chat, served by the phone and answering from Qwen3 1.7B running on it.
What it does
- A server, not a chat app
- The agents, SDKs and tools you already use call a model on your phone the way they call a cloud: OpenAI's Chat Completions, Completions and Responses, and Anthropic's Messages.
- The KV cache belongs to the server
- Clients resend the whole conversation, as both APIs expect; the server keeps the runtime's cache and the exact bytes of each reply, so follow-ups and tool loops still hit it.
- Several models, one phone
- Up to three resident, each with its own cache and endpoints, sharing one compute lane so they never halve each other's speed.
- Built to stay up
- A foreground service with the right locks, tested through forced deep Doze with the screen off. Where a phone maker's ROM can still stop it, the app names what to allow.
- A browser chat inside
- A small chat ships inside the server, so a laptop or tablet on your network can talk to the phone's models with nothing installed.
- A multiplatform core
- Everything but the runtime binding and the app shell is Kotlin Multiplatform and compiles for iOS on every build.
Measured on a phone
On a POCO X8 Pro Max (Dimensity 9500s), over Wi-Fi, with the minified release build.
| Check | Result |
|---|---|
| Official OpenAI SDK suite | 16 of 16 pass |
| Official Anthropic SDK suite | 8 of 8 pass |
| Edge-case probe (refusals, limits, disconnects) | 28 of 28 pass |
| Screen-off soak on battery, 10 minutes | 20 of 20 answered |
| Warm second turn of a 2,000-token conversation | 0.47 s instead of 8.4 s |
| Qwen3 tool loop, second request | 174 of 205 prompt tokens reused |
| XNNPACK export against llama.cpp, Snapdragon 8 Elite | 1.2 to 1.6× faster decode |
Method and raw results: the POCO X8 Pro Max report. The decode advantage of the XNNPACK export over llama.cpp was measured in OpenWeights, on whose ExecuTorch engine ExecuServe is built.
Quick start
An arm64 phone on Android 12 or later, adb on your computer, and an ExecuTorch export with its tokenizer.
From a clone of the repository, with the app installed:
tools/execuserve --model ~/models/Qwen3-1.7B-8da4w-gptq-2k.pte
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1
export OPENAI_API_KEY=<the key it prints>Then any OpenAI client:
from openai import OpenAI
client = OpenAI()
reply = client.chat.completions.create(
model="qwen3-1.7b",
messages=[{"role": "user", "content": "Hello"}],
)The script pushes the model, starts the server, forwards the port over adb, and prints the base URL and key. Anthropic SDKs, Open WebUI, the codex CLI and other apps on the same phone work too; the README has the settings for each.
The API
OpenAI's and Anthropic's shapes, error codes and streaming, checked with their official SDKs.
| POST /v1/chat/completions | Streaming, tools and tool calls, reasoning, usage with cached tokens |
| POST /v1/responses | Items, function calls, typed stream events, previous_response_id |
| POST /v1/messages | Anthropic's Messages API: tool_use, thinking blocks, its stream events |
| POST /v1/completions | A raw prompt, no template |
| GET /v1/models | Installed models with context length, residency and capabilities |
| GET /metrics | Prometheus counters, gauges and quantiles |
| /models/{id}/v1 | The model, inference and status routes scoped to one model, for hosting several at once |
Private by construction
No account, no analytics, no crash reporter and no server of ours. Prompts and replies are answered on the phone and never sent to us; the run history keeps figures, never what was asked. Every request needs a key, loopback included, and the browser chat is served under a strict Content-Security-Policy. The privacy policy says exactly what stays and what leaves.
Open source, and what comes next
Apache-2.0, on GitHub. The server core is Kotlin Multiplatform and already compiles for iOS; next are the iOS app, the other ExecuTorch backends (Vulkan, QNN, MediaTek), server-side tools, vision input, and TLS for network mode. Not affiliated with, endorsed by or sponsored by the PyTorch Foundation or Meta.