A 2B model in 10 GGUF formats on one A100.
The checkpoint is a Qwen3.5-2B tool-calling model from OpenGrad, the research repository of Experimental Intelligence. It was converted to BF16 GGUF and quantized 9 ways with llama.cpp and one importance matrix, then benchmarked on one NVIDIA A100-SXM4-80GB with every layer on the GPU. Every number is generated from the committed results. What quantization did to the model's decisions is covered in full on the OpenGrad study page. All 10 files are on Hugging Face.
Throughput per format
Against BF16, the quantized formats generate 7% to 22% faster, and process long prompts at 46% to 63% of its rate. On this GPU, quantization mainly saves memory.
Mean of 3 llama-bench repetitions; whiskers are one standard deviation. The dashed line is BF16. Prompt throughput is shown at 2,048 tokens because the 512-token runs had a standard deviation of up to 28% of their mean.
Size, bits per weight and what each format keeps
| Format | File | Bits per weight | Generation, tok/s | vs BF16 | Prompt 2,048, tok/s | Decisions kept | Quality gate |
|---|---|---|---|---|---|---|---|
| Q2_K | 0.90 GiB | 4.12 | 284.8 | +13% | 9,623 | 53.4% | fails |
| IQ3_M | 0.99 GiB | 4.50 | 283.4 | +13% | 12,420 | 77.1% | fails |
| Q3_K_M | 1.02 GiB | 4.67 | 268.8 | +7% | 10,879 | 77.1% | fails |
| IQ4_XS | 1.11 GiB | 5.08 | 306.3 | +22% | 13,288 | 88.3% | fails |
| Q4_K_S | 1.13 GiB | 5.15 | 299.2 | +19% | 12,436 | 87.9% | fails |
| Q4_K_M | 1.19 GiB | 5.42 | 292.7 | +17% | 12,200 | 87.2% | fails |
| Q5_K_M | 1.31 GiB | 6.00 | 292.9 | +17% | 12,414 | 95.0% | fails |
| Q6_Krecommended | 1.45 GiB | 6.62 | 267.7 | +7% | 11,914 | 98.0% | passes |
| Q8_0 | 1.87 GiB | 8.55 | 273.8 | +9% | 13,297 | 98.7% | passes |
| BF16 | 3.52 GiB | 16.05 | 251.1 | — | 21,073 | reference | passes |
Bits per weight is the file size over 1.88B parameters, metadata included. Decisions kept is the share of 1,277 tool-use test prompts where the format decided the same thing as BF16. The quality gate is OpenGrad's, frozen before any format existed. Q6_K is the smallest format that passes it, by one example, and OpenGrad's errata show that margin is within rerun noise. Full method in the quantization report.
The BF16 GGUF against a serving engine
These files are the deployment format. The same weights also ran as the plain BF16 checkpoint behind vLLM on an NVIDIA H200, batching up to 256 concurrent requests. This is not a like-for-like comparison: the sections above are llama.cpp on an NVIDIA A100-SXM4-80GB, and the engine, the GPU and the kernels all differ. Read it as one engine against another, not as A100 against H200.
Top: vLLM output throughput as concurrency rises on a mixed request shape. The dashed line is this page's BF16 GGUF on the A100 at a single stream, 251.1 tokens per second, drawn at one request. Bottom: decisions changed out of 1,277 confirmatory prompts, each counted against the llama.cpp BF16 GGUF.
| Concurrent | Shape | Requests/s | Output tok/s | Total tok/s |
|---|---|---|---|---|
| 1 | long | 8.2 | 139.1 | 42,977 |
| 1 | median | 5.3 | 270.5 | 4,455 |
| 1 | mixed | 8.5 | 279.6 | 407 |
| 1 | short | 6.8 | 225.2 | 328 |
| 8 | mixed | 41.7 | 1158.5 | 1,827 |
| 32 | mixed | 118.2 | 3323.5 | 5,901 |
| 64 | mixed | 208.4 | 5814.6 | 12,043 |
| 128 | mixed | 269.6 | 7794.3 | 41,524 |
| 256 | mixed | 259.9 | 8312.7 | 75,138 |
Requests per second peak at 128 concurrent and dip slightly at 256, while total tokens per second keeps climbing — the later gain is prompt tokens, not decode. One request of each shape was also run separately: long 139, median 271, mixed 280, short 225 output tokens per second. Cold start, loading the checkpoint into vLLM, was 52.6 seconds.
What the engine changed
Changing the engine, the GPU and the kernels together moved 21 of 1,277 decisions — 98% agreement, and 87% on exact output. Of those flips, 3 fall on the six prompts already known to tokenize differently between the two engines, leaving 18 attributable to engine numerics. For scale, the Q6_K quantization this page recommends changed 25. Moving between inference engines cost about as much per-example disagreement as the quantization OpenGrad spent a phase gating, which is why the comparison is reported here rather than folded into the ladder. The vLLM rerun reproduced the frozen reference within 0.011 call F1 at a parse-valid rate of 100%, so both engines were scored against a confirmed reference. Method and the per-example flips are in the H200 run report and the agreement record.
Conditions
| Engine | llama.cpp d3146f2, CUDA backend, llama-bench |
| Hardware | One NVIDIA A100-SXM4-80GB, on Modal |
| Offload | all layers on the GPU |
| Batching | 2048 batch, 512 micro-batch, 24 CPU threads |
| KV cache | f16 keys, f16 values |
| Repetitions | 3 per test |
| Run date | 2026-09-12 |
| Engine, vLLM run | vLLM 0.29.0, torch 2.13.0+cu130, CUDA 13.0 |
| Hardware, vLLM run | One NVIDIA H200, 139.8 GiB, on Modal |
| Batching, vLLM run | 5760-token context, up to 256 sequences, prefix caching off, bfloat16 |
| Run date, vLLM run | 2026-09-12 |
The throughput rows are single-stream llama-bench runs from an empty context. The vLLM run is a separate batching sweep on different hardware, so the two are not comparable row for row. Laptop, phone and CPU throughput were not measured for these files; on a phone, where memory bandwidth is the limit, the ranking can differ.