LLM inference on datacenter, laptop and phone silicon.

ServerLaptopPhone5,000 tok/s10,000 tok/sNVIDIA H200AMD MI300X11,337 tok/s8,085 tok/s11 concurrent streamH200 199 tok/sMI300X 195 tok/s44 concurrent streamsH200 761 tok/sMI300X 603 tok/s88 concurrent streamsH200 1,445 tok/sMI300X 1,137 tok/s1616 concurrent streamsH200 2,674 tok/sMI300X 2,066 tok/s3232 concurrent streamsH200 4,486 tok/sMI300X 3,443 tok/s6464 concurrent streamsH200 6,806 tok/sMI300X 4,956 tok/s128128 concurrent streamsH200 9,937 tok/sMI300X 7,022 tok/s256256 concurrent streamsH200 11,337 tok/sMI300X 8,085 tok/sAggregate tok/s over all streams. Concurrency doubles per step.Llama 3.1 8B BF16, vLLM, 1k in / 1k outApple M5Snapdragon X2 Elite14475Gemma 3 1B Q4Gemma 3 1B Q4M5 143.8 tok/sX2 Elite 74.5 tok/s2618Qwen3 8B Q4Qwen3 8B Q4M5 25.6 tok/sX2 Elite 18.2 tok/sToken generation, one stream, tok/sllama.cpp on each chip's GPUSnapdragon 8 EliteDimensity 9500s62 to 6422 to 31CPUCPU8 Elite 62 to 64 tok/s9500s 22 to 31 tok/s41.5 to 42.5~27GPUGPU8 Elite 41.5 to 42.5 tok/s9500s ~27 tok/sGemma 3 1B Q4, one stream, tok/sllama.cpp on the phone, range across runs
Server throughput is the sum over all concurrent streams; laptop and phone are one stream each. Hover a step or a bar pair for the exact numbers. Every value is from the linked repository.

Winner and margin, every lane

Bars grow toward the chip that wins, on a log scale.

Server

  • NVIDIA H200 wins
  • AMD MI300X wins

The answer flips with model size: 70B plus its KV cache no longer fits in 141 GB. 70B BF16 loads only on the MI300X.

MoE decode, 1 stream
2.72xNVIDIA H200 229 tok/s against AMD MI300X 84 tok/s, 2.72x
8B prefill-heavy, 256 streams
2.06xNVIDIA H200 1,740 tok/s against AMD MI300X 847 tok/s, 2.06x
8B balanced, 256 streams
1.4xNVIDIA H200 11,337 tok/s against AMD MI300X 8,085 tok/s, 1.4x
8B training, BF16
1.32 to 1.40xNVIDIA H200 6,707 eager, 7,822 compiled tok/s against AMD MI300X 5,102 eager, 5,603 compiled tok/s, 1.32 to 1.40x
70B FP8 balanced, 128 streams
1.14xAMD MI300X 1,873 tok/s against NVIDIA H200 1,637 tok/s, 1.14x
70B FP8 decode-heavy, 256 streams
1.08xAMD MI300X 827 tok/s against NVIDIA H200 763 tok/s, 1.08x
70B FP8 first token p99, 1k in, 8k out
62.9xAMD MI300X 35.0 s against NVIDIA H200 2,200.5 s, 62.9x
10x10x100x100xeven

Laptop

  • Apple M5 wins
  • Snapdragon X2 Elite wins

The Snapdragon wins the CPU and the NPU. Training is 19.5x because PyTorch has no Adreno backend, so the X2 trains on its CPU.

GPU prefill, Gemma 3 1B
2.92xApple M5 5,603.7 tok/s against Snapdragon X2 Elite 1,920.5 tok/s, 2.92x
GPU decode, Gemma 3 1B
1.93xApple M5 143.8 tok/s against Snapdragon X2 Elite 74.5 tok/s, 1.93x
CPU decode, Gemma 3 1B
1.1xSnapdragon X2 Elite 127.9 tok/s against Apple M5 116.3 tok/s, 1.1x
CPU prefill, Gemma 3 1B
1.9xSnapdragon X2 Elite 1,689.3 tok/s against Apple M5 889.2 tok/s, 1.9x
NPU fp16 GEMM
1.08xSnapdragon X2 Elite 15,205 GFLOP/s against Apple M5 14,137 GFLOP/s, 1.08x
Training, PyTorch
19.5xApple M5 28,060 tok/s against Snapdragon X2 Elite 1,436 tok/s, 19.5x
10x10x100x100xeven

Phone

  • Snapdragon 8 Elite wins
  • Dimensity 9500s wins

The Snapdragon wins every lane, by 1.6x on GPU decode and 18.5x on GPU prefill. The Dimensity's OpenCL is blocked, so it runs Vulkan only.

NPU int8 dense
2.9xSnapdragon 8 Elite 4,307 GOPS against Dimensity 9500s 1,100 to 1,470 GOPS, 2.9x
CPU prefill, Gemma 3 1B
2.6xSnapdragon 8 Elite 171 to 173 tok/s against Dimensity 9500s 58 to 74 tok/s, 2.6x
CPU decode, Gemma 3 1B
2.4xSnapdragon 8 Elite 62 to 64 tok/s against Dimensity 9500s 22 to 31 tok/s, 2.4x
GPU prefill, Gemma 3 1B
18.5xSnapdragon 8 Elite 717 to 723 tok/s against Dimensity 9500s ~39 tok/s, 18.5x
GPU decode, Gemma 3 1B
1.6xSnapdragon 8 Elite 41.5 to 42.5 tok/s against Dimensity 9500s ~27 tok/s, 1.6x
GPU matmul, f32
7.1xSnapdragon 8 Elite 479 GFLOP/s against Dimensity 9500s 67 GFLOP/s, 7.1x
GPU training step
8.1xSnapdragon 8 Elite 403 GFLOP/s against Dimensity 9500s ~50 GFLOP/s, 8.1x
10x10x100x100xeven

First token and per token on five phone chips

Five small models on llama.cpp and ExecuTorch, measured inside the OpenWeights app. Medians over 60 to 90 prompts per cell. Interactive version.

Gemma 3 1B

First token, s
0.31310
Per token, ms
2050100200
D9400
8 Gen 3
8 Elite
Tensor G5
Exynos 2400

LFM2.5 1.2B

First token, s
0.31310
Per token, ms
2050100200
D9400
8 Gen 3
8 Elite
Tensor G5
Exynos 2400

Qwen3 1.7B

First token, s
0.31310
Per token, ms
2050100200
D9400
8 Gen 3
8 Elite
Tensor G5
Exynos 2400

Llama 3.2 3B

First token, s
0.31310
Per token, ms
2050100200
D9400
8 Gen 3
8 Elite
Tensor G5
Exynos 2400

SmolLM3 3B

First token, s
0.31310
Per token, ms
2050100200
D9400
8 Gen 3
8 Elite
Tensor G5
Exynos 2400
  • llama.cpp
  • ExecuTorch

Both axes are log scales, shared across the five panels. Hover or tap a mark for its number; the rate in parentheses is the median prefill or decode tokens per second.

Conditions are not matched. The Dimensity 9400 ran in hand under a fan; the other four ran racked in cloud test labs, where phones throttle within about 40 seconds. Single pass, 2026-09-03 and 04.

GPU against ASIC

At 32 concurrent streams, one Inferentia2 cost 2.4x as much per output token as one MI300X. Trainium1 fine-tuned the same 8B model at 68.3% MFU. The comparison, with the training, serving and speculative decoding figures.

Server

AMD MI300X against NVIDIA H200 in inference and training, plus Trainium, Inferentia and TPUs.

RepositoryResult
MI300X-vs-H200
One GPU each, serving and training, three shapes, eight concurrency points
MI300X 1.14x on 70B FP8; H200 1.40x on 8B serving and 1.32 to 1.40x on training
qwen3.8-27b-mi300x
Qwen3.8-27B served from one MI300X with vLLM behind an authenticated endpoint
OpenAI-compatible endpoint
torchneuronx
Llama 3.1 8B LoRA on Trainium1, served by vLLM on Inferentia2
trn1 to inf2
serverless-inference
Scale-to-zero inference on RunPod; model and GPU chosen by shootout
RunPod, scale to zero
compute-visualizer
Roofline and five-way bottleneck analysis for H100 training and inference
H100 roofline, five bottlenecks
will-it-asic
Will this model fit a TPU, Trainium, Inferentia, Gaudi or GPU?
TPU, Trainium, Inferentia, Gaudi, GPU
pytorch-for-asics
De-mystifying PyTorch for ASICs, PyTorch Conference Europe 2026
conference talk
xla-agentic-development
Skills for coding agents on TPUs: Pallas, XProf, XLA lowering
Claude Code and Codex plugin

Laptop

Apple M5 against Snapdragon X2 Elite: same llama.cpp release, byte-identical weights, each chip's own GPU backend.

RepositoryResult
snapdragon-vs-m5
37 tests each: CPU, GPU and NPU inference, training, a 10-minute sustained loop, perplexity
M5 1.93x on GPU decode; X2 1.10x on CPU
mlx-models
MLP, CNN and ViT trained from scratch on an M5 Air, then Whisper, CLIP, SigLIP
MLX 0.32 on 24 GB unified memory
mlx-agentic-development
Does an MLX skills kit help a coding agent? Pre-registered, placebo arm, 250 runs
null result, p = 0.69

Phone

Snapdragon 8 Elite against Dimensity 9500s, measured on the phones themselves, no root.

RepositoryResult
snapdragon-vs-mediatek
NPU, GPU and CPU inference and training, int8 and fp16, on device
4,307 vs 1,100 to 1,470 GOPS int8 on the NPU
poco-phone-ai-training
Is the Dimensity 9500s NPU reachable without root? Through NeuroPilot, yes
1.1 to 1.5 TOPS int8
openweights
Hugging Face open weights on Android, native Kotlin and llama.cpp, no account
on the Play Store

Have a claim worth testing?

alpha@experimentalmachines.org