ExecuTorch exports: what is published, and what is still missing.
execupack exports small dense open-weight LLMs, 4B parameters or fewer, from Hugging Face to ExecuTorch .pte files for the openweights Android app, and publishes them under experimentalmachines. This page sets every published file against what ExecuTorch 1.4.0, the version the app's runtime ships, can export at all, family by family and accelerator by accelerator.
- models published
- 16
- models published
- builds (window × chip)
- 179
- builds (window × chip)
- family × accelerator pairs ExecuTorch supports, not yet published
- 16
- family × accelerator pairs ExecuTorch supports, not yet published
- checkpoints ExecuTorch names, no repo yet
- 13
- checkpoints ExecuTorch names, no repo yet
By family and accelerator
Each cell is one model family on one accelerator. "Not yet" means ExecuTorch 1.4.0 can export it and nothing is on Hugging Face; "not in 1.4.0" means there is no export path to follow.
| Family | XNNPACKCPU, any Android phone | VulkanGPU, any Vulkan phone | Qualcomm QNNSnapdragon NPU, per chip | MediaTekDimensity NPU, per chip | Samsung ExynosExynos NPU, per chip |
|---|---|---|---|---|---|
| Qwen30.6B, 1.7B, 4B | publishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507 | publishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507 | not yetQwen3-0.6B, Qwen3-1.7B | publishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507MediaTek's stock scripts build RoPE at base 10000; execupack's mediatek-rope-theta.patch makes the fp32 graph match Hugging Face exactly | not in 1.4.0delegate exists, no LLM example |
| Qwen2.50.5B, 1.5B, 3B | publishedQwen2.5-0.5B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-Math-1.5B-Instruct | publishedQwen2.5-0.5B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-Math-1.5B-Instruct | not yetQwen2.5-0.5B, Qwen2.5-1.5B; base checkpoints only | refusedthe runner's -100 attention mask leaks through Qwen2.5's scores: KL 0.013 and 98% top-1 against Hugging Face before any quantization | not in 1.4.0delegate exists, no LLM example |
| Llama 3.21B, 3B | publishedLlama-3.2-1B-Instruct, Llama-3.2-3B-Instruct | publishedLlama-3.2-1B-Instruct, Llama-3.2-3B-Instruct | not yetLlama-3.2-1B-Instruct, Llama-3.2-3B-Instruct | publishedLlama-3.2-1B-Instruct, Llama-3.2-3B-InstructMediaTek's Llama model lacks rope_theta and Llama 3 RoPE scaling; execupack's mediatek-rope-theta.patch adds both | not in 1.4.0delegate exists, no LLM example |
| SmolLM2135M, 360M, 1.7B | publishedSmolLM2-135M-Instruct, SmolLM2-360M-Instruct | publishedSmolLM2-135M-Instruct, SmolLM2-360M-Instruct | not yetSmolLM2-135M-Instruct | publishedSmolLM2-135M-Instruct, SmolLM2-360M-Instructrope_theta from execupack's patch, and the fast tokenizer instead of Llama's SentencePiece default | not in 1.4.0delegate exists, no LLM example |
| LFM2 / LFM2.5350M, 700M, 1.2B, 2.6B | publishedLFM2.5-1.2B-Instruct, LFM2.5-1.2B-Instruct-heretic, LFM2.5-2.6B, LFM2.5-2.6B-heretic | not in 1.4.0no Vulkan kernel for the short convolution; a file that lowers anyway segfaults at the first prefill | not in 1.4.0not in Qualcomm's SUPPORTED_LLM_MODELS registry | publishedLFM2.5-1.2B-Instruct, LFM2.5-1.2B-Instruct-heretic, LFM2.5-2.6B, LFM2.5-2.6B-hereticno LFM2 model in examples/mediatek; execupack's mediatek-lfm2.patch adds one | not in 1.4.0delegate exists, no LLM example |
| Phi-4-mini3.8B | not yet | not yet | not yetPhi-4-mini-instruct | not yetmodel_type phi4 | not in 1.4.0delegate exists, no LLM example |
| Gemma 31B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetgemma-3-1b-it | not yetmodel_type gemma3; HF configs say gemma3_text | not in 1.4.0delegate exists, no LLM example |
| Gemma 22B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetgemma-2-2b-it | not yetmodel_type gemma2 | not in 1.4.0delegate exists, no LLM example |
| SmolLM33B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetSmolLM3-3B | not in 1.4.0no SmolLM3 model in examples/mediatek | not in 1.4.0delegate exists, no LLM example |
| Gemma2B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetgemma-2b-it | not in 1.4.0examples/mediatek resolves gemma2 and gemma3, not gemma | not in 1.4.0delegate exists, no LLM example |
| GLM-Edge1.5B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetglm-edge-1.5b-chat | not in 1.4.0no GLM model in examples/mediatek | not in 1.4.0delegate exists, no LLM example |
| Granite 3.32B | not in 1.4.0not in export_llm's ModelType list | not in 1.4.0not in export_llm's ModelType list | not yetgranite-3.3-2b-instruct | not in 1.4.0no Granite model in examples/mediatek | not in 1.4.0delegate exists, no LLM example |
| Qwen3.50.8B, 2B, 4B | blockedthe openweights app refuses names containing qwen35, so execupack does not export them | blockedthe openweights app refuses names containing qwen35, so execupack does not export them | not in 1.4.0not in Qualcomm's SUPPORTED_LLM_MODELS registry | not in 1.4.0no Qwen3.5 model in examples/mediatek | not in 1.4.0delegate exists, no LLM example |
XNNPACK and Vulkan files run on any phone; NPU files are compiled for one chip and load only on it, so a published QNN or MediaTek cell covers the chips listed in the next table, not every Snapdragon or Dimensity. Vulkan uses the same export_llm recipe as XNNPACK with the GPU delegate, and ExecuTorch accepts it for every model class it lists. One Vulkan file has been run on a phone: Qwen3-0.6B at 2k on a Dimensity 9400 (Mali GPU), with ExecuTorch 1.4.0's own runner built with the Vulkan delegate. It answered correctly at 18 tokens per second, against 52 for the XNNPACK file on the same phone. The executorch-android 1.4.0 library that the openweights app ships registers XNNPACK only, so the app cannot load Vulkan files until it moves to executorch-android-vulkan 1.4.0, which registers both. Qualcomm's scripts export a fixed list of checkpoints, each with its own quantization recipe, so a fine-tune of a listed model is not covered.
By model
Every published repo, with the context windows present for each accelerator, followed by the checkpoints ExecuTorch names that have no repo yet. 5 accelerator builds are possible for published models and missing.
| Model | XNNPACK | Vulkan | Qualcomm QNN | MediaTek |
|---|---|---|---|---|
| Qwen3-0.6BQwen3 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | not yetcompiled per chip | 4k 8kMT6991, A16W8 |
| Qwen3-1.7BQwen3 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | not yetcompiled per chip | 2k 4k 8kMT6991, A16W8 |
| Qwen3-4BQwen3 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | — | 2k 4kMT6991, A16W8 |
| Qwen3-4B-Instruct-2507Qwen3 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | — | 2k 4kMT6991, A16W8 |
| Qwen2.5-0.5B-InstructQwen2.5 · updated 2026-10-03 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | — | — |
| Qwen2.5-1.5B-InstructQwen2.5 · updated 2026-10-03 | 2k 4k 8k 16k8da4w | 2k 4k 8k 16k 32k8da4w | — | — |
| Qwen2.5-3B-InstructQwen2.5 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | — | — |
| Qwen2.5-Math-1.5B-InstructQwen2.5 · updated 2026-10-03 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | — | — |
| Qwen2.5-0.5BQwen2.5 · no repo yet | not yet | not yet | not yetcompiled per chip | — |
| Qwen2.5-1.5BQwen2.5 · no repo yet | not yet | not yet | not yetcompiled per chip | — |
| Llama-3.2-1B-InstructLlama 3.2 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | not yetcompiled per chip | 2k 4k 8k 16kMT6991, A16W8 |
| Llama-3.2-3B-InstructLlama 3.2 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | not yetcompiled per chip | 2k 4k 8kMT6991, A16W8 |
| SmolLM2-135M-InstructSmolLM2 · updated 2026-10-03 | 2k 4k 8k 16k 32k8da4w GPTQ | 2k 4k 8k 16k 32k8da4w | not yetcompiled per chip | 2k 4k 8k 16kMT6991, A16W8 |
| SmolLM2-360M-InstructSmolLM2 · updated 2026-10-04 | 2k 4k 8k 16k 32kfp32 linears | 2k 4k 8k 16k 32k8da4w | — | 2k 4k 8k 16kMT6991, A16W8 |
| LFM2.5-1.2B-InstructLFM2 / LFM2.5 · updated 2026-09-24 | 2k 4k 8k 16k 32k8da4w GPTQ | — | — | 512 2k 4k 8kMT6991, A16W8 |
| LFM2.5-1.2B-Instruct-hereticLFM2 / LFM2.5 · updated 2026-10-03 | 2k 4k 8k 16k 32k8da4w GPTQ | — | — | 512 2k 4k 8kMT6991, A16W8 |
| LFM2.5-2.6BLFM2 / LFM2.5 · updated 2026-10-04 | 2k 4k 8k 16k 32k8da4w GPTQ | — | — | 512MT6991, A16W4 2k 4k 8kMT6991, A16W8 |
| LFM2.5-2.6B-hereticLFM2 / LFM2.5 · updated 2026-10-03 | 2k 4k 8k 16k 32k8da4w GPTQ | — | — | 512 2k 4k 8kMT6991, A16W8 |
| LFM2-1.2BLFM2 / LFM2.5 · no repo yet | not yet | — | — | not yetcompiled per chip |
| LFM2-350MLFM2 / LFM2.5 · no repo yet | not yet | — | — | not yetcompiled per chip |
| LFM2-700MLFM2 / LFM2.5 · no repo yet | not yet | — | — | not yetcompiled per chip |
| LFM2.5-350MLFM2 / LFM2.5 · no repo yet | not yet | — | — | not yetcompiled per chip |
| Phi-4-mini-instructPhi-4-mini · no repo yet | not yet | not yet | not yetcompiled per chip | not yetcompiled per chip |
| gemma-3-1b-itGemma 3 · no repo yet | — | — | not yetcompiled per chip | not yetcompiled per chip |
| gemma-2-2b-itGemma 2 · no repo yet | — | — | not yetcompiled per chip | not yetcompiled per chip |
| SmolLM3-3BSmolLM3 · no repo yet | — | — | not yetcompiled per chip | — |
| gemma-2b-itGemma · no repo yet | — | — | not yetcompiled per chip | — |
| glm-edge-1.5b-chatGLM-Edge · no repo yet | — | — | not yetcompiled per chip | — |
| granite-3.3-2b-instructGranite 3.3 · no repo yet | — | — | not yetcompiled per chip | — |
Windows are context lengths in tokens (2k = 2,048). XNNPACK files marked 8da4w GPTQ or fp32 linears passed execupack's decision gate against the fp32 model before publishing. Qwen2.5-1.5B-Instruct's are the older round-to-nearest build, which fails that gate, and no int4 build of it passes; its fp32 file would be 6.2 GB. A dash means ExecuTorch 1.4.0 has no path for that model on that accelerator. Every repo also carries the tokenizer at its root and a config.json per backend folder that the app reads. The file list is generated from the Hugging Face API; the newest change it saw was on 2026-10-04.
Chips ExecuTorch can export to
XNNPACK and Vulkan files run on any Android phone. NPU files are compiled for one chip and load only on that chip, so every chip is its own export. These are the chips each NPU delegate in ExecuTorch 1.4.0 can compile for.
Qualcomm QNN
QcomChipset lists 20 chips. execupack compiles with QAIRT 2.37, the SDK the executorch 1.4.0 wheel downloads, which reaches HTP V79; the V81 chips need QAIRT 2.42 or newer. Source: backends/qualcomm/serialization/qc_schema.py.
| Chip | Product | Architecture | Status |
|---|---|---|---|
SM8750 | Snapdragon 8 Elite | V79 | execupack targets itexecupack's QNN chip (qnn.socs) |
SM8650 | Snapdragon 8 Gen 3 | V75 | can export |
SM8550 | Snapdragon 8 Gen 2 | V73 | can export |
SM8475 | Snapdragon 8+ Gen 1 | V69 | can exportno block 4-bit (LPBQ) or 16-bit matmul input below V73 |
SM8450 | Snapdragon 8 Gen 1 | V69 | can exportno block 4-bit (LPBQ) or 16-bit matmul input below V73 |
SM8350 | Snapdragon 888 | V68 | can exportV68: the registry's default LLM recipes need 8-bit fallbacks |
SM8850 | Snapdragon 8 Elite Gen 5 | V81 | needs a newer SDK |
SM8845 | Snapdragon 8 Gen 5 | V81 | needs a newer SDK |
QCM6490 | IoT | V68 | can export |
SA8295 | Automotive | V68 | can export |
SA8255 | Automotive | V73 | can export |
QCS9100 | Automotive / industrial | V73 | can export |
SSG2115P | XR / glasses | V73 | can export |
SSG2125P | XR / glasses | V73 | can export |
SXR1230P | XR | V73 | can export |
SXR2230P | XR (Meta Quest 3) | V69 | can export |
SXR2330P | XR | V79 | can export |
SA8797 | Automotive | V81 | needs a newer SDK |
SAR2230P | XR | V81 | needs a newer SDK |
SW6100 | Wearable | V81 | needs a newer SDK |
MediaTek NeuroPilot
The delegate accepts three platforms, but the LLM export scripts in examples/mediatek offer only DX3 and DX4 (--platform), so Dimensity 9500 has a delegate and no LLM path without patching them. Source: backends/mediatek/preprocess.py.
| Chip | Product | Architecture | Status |
|---|---|---|---|
MT6991 | Dimensity 9400 | DX4 | execupack targets itexecupack's MediaTek chip (mtk.socs) |
MT6989 | Dimensity 9300 | DX3 | can export |
MT6993 | Dimensity 9500 | — | delegate onlyin SUPPORTED_PLATFORM_CONFIGS, not in the LLM scripts' --platform choices |
Samsung Exynos
The EnnBackend delegate supports two chipsets, and ExecuTorch 1.4.0 has no LLM example for either. Source: backends/samsung/README.md.
| Chip | Product | Architecture | Status |
|---|---|---|---|
E9955 | Exynos 2500 | — | no LLM path |
E9965 | Exynos 2600 | — | no LLM path |
Sources
What ExecuTorch can export is read from its source at the v1.4.0 tag, which defines the registries below. execupack pins the same version because a newer exporter can emit methods the app's runtime lacks. Decisions and measurements behind each backend are in execupack's plan.
extension/llm/export/config/llm_config.py | ModelType: the model classes export_llm builds, for XNNPACK and Vulkan |
examples/qualcomm/oss_scripts/llama/__init__.py | SUPPORTED_LLM_MODELS: the Qualcomm registry, one entry per checkpoint |
examples/mediatek/aot_utils/llm_utils/utils.py | resolve_model_classes: the config model_types MediaTek's scripts build |
backends/samsung | the Exynos delegate; examples/samsung has CNN examples only |
Left out on purpose
The same registries list these, and execupack does not export them:
| stories110M, stories260K | toy checkpoints for testing |
| Codegen2 1B | code completion, not chat |
| Llama 2, Llama 3, Llama 3.1, Qwen2.5-Coder 32B | above 4B |
| Qwen3.5 MoE | mixture of experts |
| Llama 3.2 Vision, InternVL3, SmolVLM, Granite Speech, Gemma 4 | multimodal; the app runs text models |
iOS accelerators (Core ML, MPS) are outside this page: execupack exports for Android only, and iOS is deferred.