ExecuTorch exports: what is published, and what is still missing.

execupack exports small dense open-weight LLMs, 4B parameters or fewer, from Hugging Face to ExecuTorch .pte files for the openweights Android app, and publishes them under experimentalmachines. This page sets every published file against what ExecuTorch 1.4.0, the version the app's runtime ships, can export at all, family by family and accelerator by accelerator.

models published
16
models published
builds (window × chip)
179
builds (window × chip)
family × accelerator pairs ExecuTorch supports, not yet published
16
family × accelerator pairs ExecuTorch supports, not yet published
checkpoints ExecuTorch names, no repo yet
13
checkpoints ExecuTorch names, no repo yet

By family and accelerator

Each cell is one model family on one accelerator. "Not yet" means ExecuTorch 1.4.0 can export it and nothing is on Hugging Face; "not in 1.4.0" means there is no export path to follow.

FamilyXNNPACKCPU, any Android phoneVulkanGPU, any Vulkan phoneQualcomm QNNSnapdragon NPU, per chipMediaTekDimensity NPU, per chipSamsung ExynosExynos NPU, per chip
Qwen30.6B, 1.7B, 4BpublishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507publishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507not yetQwen3-0.6B, Qwen3-1.7BpublishedQwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-4B-Instruct-2507MediaTek's stock scripts build RoPE at base 10000; execupack's mediatek-rope-theta.patch makes the fp32 graph match Hugging Face exactlynot in 1.4.0delegate exists, no LLM example
Qwen2.50.5B, 1.5B, 3BpublishedQwen2.5-0.5B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-Math-1.5B-InstructpublishedQwen2.5-0.5B-Instruct, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-Math-1.5B-Instructnot yetQwen2.5-0.5B, Qwen2.5-1.5B; base checkpoints onlyrefusedthe runner's -100 attention mask leaks through Qwen2.5's scores: KL 0.013 and 98% top-1 against Hugging Face before any quantizationnot in 1.4.0delegate exists, no LLM example
Llama 3.21B, 3BpublishedLlama-3.2-1B-Instruct, Llama-3.2-3B-InstructpublishedLlama-3.2-1B-Instruct, Llama-3.2-3B-Instructnot yetLlama-3.2-1B-Instruct, Llama-3.2-3B-InstructpublishedLlama-3.2-1B-Instruct, Llama-3.2-3B-InstructMediaTek's Llama model lacks rope_theta and Llama 3 RoPE scaling; execupack's mediatek-rope-theta.patch adds bothnot in 1.4.0delegate exists, no LLM example
SmolLM2135M, 360M, 1.7BpublishedSmolLM2-135M-Instruct, SmolLM2-360M-InstructpublishedSmolLM2-135M-Instruct, SmolLM2-360M-Instructnot yetSmolLM2-135M-InstructpublishedSmolLM2-135M-Instruct, SmolLM2-360M-Instructrope_theta from execupack's patch, and the fast tokenizer instead of Llama's SentencePiece defaultnot in 1.4.0delegate exists, no LLM example
LFM2 / LFM2.5350M, 700M, 1.2B, 2.6BpublishedLFM2.5-1.2B-Instruct, LFM2.5-1.2B-Instruct-heretic, LFM2.5-2.6B, LFM2.5-2.6B-hereticnot in 1.4.0no Vulkan kernel for the short convolution; a file that lowers anyway segfaults at the first prefillnot in 1.4.0not in Qualcomm's SUPPORTED_LLM_MODELS registrypublishedLFM2.5-1.2B-Instruct, LFM2.5-1.2B-Instruct-heretic, LFM2.5-2.6B, LFM2.5-2.6B-hereticno LFM2 model in examples/mediatek; execupack's mediatek-lfm2.patch adds onenot in 1.4.0delegate exists, no LLM example
Phi-4-mini3.8Bnot yetnot yetnot yetPhi-4-mini-instructnot yetmodel_type phi4not in 1.4.0delegate exists, no LLM example
Gemma 31Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetgemma-3-1b-itnot yetmodel_type gemma3; HF configs say gemma3_textnot in 1.4.0delegate exists, no LLM example
Gemma 22Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetgemma-2-2b-itnot yetmodel_type gemma2not in 1.4.0delegate exists, no LLM example
SmolLM33Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetSmolLM3-3Bnot in 1.4.0no SmolLM3 model in examples/mediateknot in 1.4.0delegate exists, no LLM example
Gemma2Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetgemma-2b-itnot in 1.4.0examples/mediatek resolves gemma2 and gemma3, not gemmanot in 1.4.0delegate exists, no LLM example
GLM-Edge1.5Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetglm-edge-1.5b-chatnot in 1.4.0no GLM model in examples/mediateknot in 1.4.0delegate exists, no LLM example
Granite 3.32Bnot in 1.4.0not in export_llm's ModelType listnot in 1.4.0not in export_llm's ModelType listnot yetgranite-3.3-2b-instructnot in 1.4.0no Granite model in examples/mediateknot in 1.4.0delegate exists, no LLM example
Qwen3.50.8B, 2B, 4Bblockedthe openweights app refuses names containing qwen35, so execupack does not export themblockedthe openweights app refuses names containing qwen35, so execupack does not export themnot in 1.4.0not in Qualcomm's SUPPORTED_LLM_MODELS registrynot in 1.4.0no Qwen3.5 model in examples/mediateknot in 1.4.0delegate exists, no LLM example

XNNPACK and Vulkan files run on any phone; NPU files are compiled for one chip and load only on it, so a published QNN or MediaTek cell covers the chips listed in the next table, not every Snapdragon or Dimensity. Vulkan uses the same export_llm recipe as XNNPACK with the GPU delegate, and ExecuTorch accepts it for every model class it lists. One Vulkan file has been run on a phone: Qwen3-0.6B at 2k on a Dimensity 9400 (Mali GPU), with ExecuTorch 1.4.0's own runner built with the Vulkan delegate. It answered correctly at 18 tokens per second, against 52 for the XNNPACK file on the same phone. The executorch-android 1.4.0 library that the openweights app ships registers XNNPACK only, so the app cannot load Vulkan files until it moves to executorch-android-vulkan 1.4.0, which registers both. Qualcomm's scripts export a fixed list of checkpoints, each with its own quantization recipe, so a fine-tune of a listed model is not covered.

By model

Every published repo, with the context windows present for each accelerator, followed by the checkpoints ExecuTorch names that have no repo yet. 5 accelerator builds are possible for published models and missing.

ModelXNNPACKVulkanQualcomm QNNMediaTek
Qwen3-0.6BQwen3 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
not yetcompiled per chip
4k 8kMT6991, A16W8
Qwen3-1.7BQwen3 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
not yetcompiled per chip
2k 4k 8kMT6991, A16W8
Qwen3-4BQwen3 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
—
2k 4kMT6991, A16W8
Qwen3-4B-Instruct-2507Qwen3 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
—
2k 4kMT6991, A16W8
Qwen2.5-0.5B-InstructQwen2.5 · updated 2026-10-03
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
——
Qwen2.5-1.5B-InstructQwen2.5 · updated 2026-10-03
2k 4k 8k 16k8da4w
2k 4k 8k 16k 32k8da4w
——
Qwen2.5-3B-InstructQwen2.5 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
——
Qwen2.5-Math-1.5B-InstructQwen2.5 · updated 2026-10-03
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
——
Qwen2.5-0.5BQwen2.5 · no repo yetnot yetnot yetnot yetcompiled per chip—
Qwen2.5-1.5BQwen2.5 · no repo yetnot yetnot yetnot yetcompiled per chip—
Llama-3.2-1B-InstructLlama 3.2 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
not yetcompiled per chip
2k 4k 8k 16kMT6991, A16W8
Llama-3.2-3B-InstructLlama 3.2 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
not yetcompiled per chip
2k 4k 8kMT6991, A16W8
SmolLM2-135M-InstructSmolLM2 · updated 2026-10-03
2k 4k 8k 16k 32k8da4w GPTQ
2k 4k 8k 16k 32k8da4w
not yetcompiled per chip
2k 4k 8k 16kMT6991, A16W8
SmolLM2-360M-InstructSmolLM2 · updated 2026-10-04
2k 4k 8k 16k 32kfp32 linears
2k 4k 8k 16k 32k8da4w
—
2k 4k 8k 16kMT6991, A16W8
LFM2.5-1.2B-InstructLFM2 / LFM2.5 · updated 2026-09-24
2k 4k 8k 16k 32k8da4w GPTQ
——
512 2k 4k 8kMT6991, A16W8
LFM2.5-1.2B-Instruct-hereticLFM2 / LFM2.5 · updated 2026-10-03
2k 4k 8k 16k 32k8da4w GPTQ
——
512 2k 4k 8kMT6991, A16W8
LFM2.5-2.6BLFM2 / LFM2.5 · updated 2026-10-04
2k 4k 8k 16k 32k8da4w GPTQ
——
512MT6991, A16W4
2k 4k 8kMT6991, A16W8
LFM2.5-2.6B-hereticLFM2 / LFM2.5 · updated 2026-10-03
2k 4k 8k 16k 32k8da4w GPTQ
——
512 2k 4k 8kMT6991, A16W8
LFM2-1.2BLFM2 / LFM2.5 · no repo yetnot yet——not yetcompiled per chip
LFM2-350MLFM2 / LFM2.5 · no repo yetnot yet——not yetcompiled per chip
LFM2-700MLFM2 / LFM2.5 · no repo yetnot yet——not yetcompiled per chip
LFM2.5-350MLFM2 / LFM2.5 · no repo yetnot yet——not yetcompiled per chip
Phi-4-mini-instructPhi-4-mini · no repo yetnot yetnot yetnot yetcompiled per chipnot yetcompiled per chip
gemma-3-1b-itGemma 3 · no repo yet——not yetcompiled per chipnot yetcompiled per chip
gemma-2-2b-itGemma 2 · no repo yet——not yetcompiled per chipnot yetcompiled per chip
SmolLM3-3BSmolLM3 · no repo yet——not yetcompiled per chip—
gemma-2b-itGemma · no repo yet——not yetcompiled per chip—
glm-edge-1.5b-chatGLM-Edge · no repo yet——not yetcompiled per chip—
granite-3.3-2b-instructGranite 3.3 · no repo yet——not yetcompiled per chip—

Windows are context lengths in tokens (2k = 2,048). XNNPACK files marked 8da4w GPTQ or fp32 linears passed execupack's decision gate against the fp32 model before publishing. Qwen2.5-1.5B-Instruct's are the older round-to-nearest build, which fails that gate, and no int4 build of it passes; its fp32 file would be 6.2 GB. A dash means ExecuTorch 1.4.0 has no path for that model on that accelerator. Every repo also carries the tokenizer at its root and a config.json per backend folder that the app reads. The file list is generated from the Hugging Face API; the newest change it saw was on 2026-10-04.

Chips ExecuTorch can export to

XNNPACK and Vulkan files run on any Android phone. NPU files are compiled for one chip and load only on that chip, so every chip is its own export. These are the chips each NPU delegate in ExecuTorch 1.4.0 can compile for.

Qualcomm QNN

QcomChipset lists 20 chips. execupack compiles with QAIRT 2.37, the SDK the executorch 1.4.0 wheel downloads, which reaches HTP V79; the V81 chips need QAIRT 2.42 or newer. Source: backends/qualcomm/serialization/qc_schema.py.

ChipProductArchitectureStatus
SM8750Snapdragon 8 EliteV79execupack targets itexecupack's QNN chip (qnn.socs)
SM8650Snapdragon 8 Gen 3V75can export
SM8550Snapdragon 8 Gen 2V73can export
SM8475Snapdragon 8+ Gen 1V69can exportno block 4-bit (LPBQ) or 16-bit matmul input below V73
SM8450Snapdragon 8 Gen 1V69can exportno block 4-bit (LPBQ) or 16-bit matmul input below V73
SM8350Snapdragon 888V68can exportV68: the registry's default LLM recipes need 8-bit fallbacks
SM8850Snapdragon 8 Elite Gen 5V81needs a newer SDK
SM8845Snapdragon 8 Gen 5V81needs a newer SDK
QCM6490IoTV68can export
SA8295AutomotiveV68can export
SA8255AutomotiveV73can export
QCS9100Automotive / industrialV73can export
SSG2115PXR / glassesV73can export
SSG2125PXR / glassesV73can export
SXR1230PXRV73can export
SXR2230PXR (Meta Quest 3)V69can export
SXR2330PXRV79can export
SA8797AutomotiveV81needs a newer SDK
SAR2230PXRV81needs a newer SDK
SW6100WearableV81needs a newer SDK

MediaTek NeuroPilot

The delegate accepts three platforms, but the LLM export scripts in examples/mediatek offer only DX3 and DX4 (--platform), so Dimensity 9500 has a delegate and no LLM path without patching them. Source: backends/mediatek/preprocess.py.

ChipProductArchitectureStatus
MT6991Dimensity 9400DX4execupack targets itexecupack's MediaTek chip (mtk.socs)
MT6989Dimensity 9300DX3can export
MT6993Dimensity 9500—delegate onlyin SUPPORTED_PLATFORM_CONFIGS, not in the LLM scripts' --platform choices

Samsung Exynos

The EnnBackend delegate supports two chipsets, and ExecuTorch 1.4.0 has no LLM example for either. Source: backends/samsung/README.md.

ChipProductArchitectureStatus
E9955Exynos 2500—no LLM path
E9965Exynos 2600—no LLM path

Sources

What ExecuTorch can export is read from its source at the v1.4.0 tag, which defines the registries below. execupack pins the same version because a newer exporter can emit methods the app's runtime lacks. Decisions and measurements behind each backend are in execupack's plan.

extension/llm/export/config/llm_config.pyModelType: the model classes export_llm builds, for XNNPACK and Vulkan
examples/qualcomm/oss_scripts/llama/__init__.pySUPPORTED_LLM_MODELS: the Qualcomm registry, one entry per checkpoint
examples/mediatek/aot_utils/llm_utils/utils.pyresolve_model_classes: the config model_types MediaTek's scripts build
backends/samsungthe Exynos delegate; examples/samsung has CNN examples only

Left out on purpose

The same registries list these, and execupack does not export them:

stories110M, stories260Ktoy checkpoints for testing
Codegen2 1Bcode completion, not chat
Llama 2, Llama 3, Llama 3.1, Qwen2.5-Coder 32Babove 4B
Qwen3.5 MoEmixture of experts
Llama 3.2 Vision, InternVL3, SmolVLM, Granite Speech, Gemma 4multimodal; the app runs text models

iOS accelerators (Core ML, MPS) are outside this page: execupack exports for Android only, and iOS is deferred.