A 2B model in 10 GGUF formats on one A100.

The checkpoint is a Qwen3.5-2B tool-calling model from OpenGrad, the research repository of Experimental Intelligence. It was converted to BF16 GGUF and quantized 9 ways with llama.cpp and one importance matrix, then benchmarked on one NVIDIA A100-SXM4-80GB with every layer on the GPU. Every number is generated from the committed results. What quantization did to the model's decisions is covered in full on the OpenGrad study page. All 10 files are on Hugging Face.

Throughput per format

Against BF16, the quantized formats generate 7% to 22% faster, and process long prompts at 46% to 63% of its rate. On this GPU, quantization mainly saves memory.

Generation, 128 tokenstokens per second0100200300400284.8283.4268.8306.3299.2292.7292.9267.7273.8251.1Prompt, 2,048 tokenstokens per second010,00020,0009,62312,42010,87913,28812,43612,20012,41411,91413,29721,073Q2_K 0.90 GiBIQ3_M 0.99 GiBQ3_K_M 1.02 GiBIQ4_XS 1.11 GiBQ4_K_S 1.13 GiBQ4_K_M 1.19 GiBQ5_K_M 1.31 GiBQ6_K 1.45 GiBQ8_0 1.87 GiBBF16 3.52 GiB

Mean of 3 llama-bench repetitions; whiskers are one standard deviation. The dashed line is BF16. Prompt throughput is shown at 2,048 tokens because the 512-token runs had a standard deviation of up to 28% of their mean.

Size, bits per weight and what each format keeps

FormatFileBits per weightGeneration, tok/svs BF16Prompt 2,048, tok/sDecisions keptQuality gate
Q2_K0.90 GiB4.12284.8+13%9,62353.4%fails
IQ3_M0.99 GiB4.50283.4+13%12,42077.1%fails
Q3_K_M1.02 GiB4.67268.8+7%10,87977.1%fails
IQ4_XS1.11 GiB5.08306.3+22%13,28888.3%fails
Q4_K_S1.13 GiB5.15299.2+19%12,43687.9%fails
Q4_K_M1.19 GiB5.42292.7+17%12,20087.2%fails
Q5_K_M1.31 GiB6.00292.9+17%12,41495.0%fails
Q6_Krecommended1.45 GiB6.62267.7+7%11,91498.0%passes
Q8_01.87 GiB8.55273.8+9%13,29798.7%passes
BF163.52 GiB16.05251.1—21,073referencepasses

Bits per weight is the file size over 1.88B parameters, metadata included. Decisions kept is the share of 1,277 tool-use test prompts where the format decided the same thing as BF16. The quality gate is OpenGrad's, frozen before any format existed. Q6_K is the smallest format that passes it, by one example, and OpenGrad's errata show that margin is within rerun noise. Full method in the quantization report.

The BF16 GGUF against a serving engine

These files are the deployment format. The same weights also ran as the plain BF16 checkpoint behind vLLM on an NVIDIA H200, batching up to 256 concurrent requests. This is not a like-for-like comparison: the sections above are llama.cpp on an NVIDIA A100-SXM4-80GB, and the engine, the GPU and the kernels all differ. Read it as one engine against another, not as A100 against H200.

vLLM on NVIDIA H200, mixed shapeoutput tokens per second, one request to 256 concurrent02,0004,0006,0008,000llama.cpp, A100, one stream: 251.1279.61,158.53,323.55,814.67,794.38,312.71 request8 concurrent32 concurrent64 concurrent128 concurrent256 concurrentDecisions changed, of 1,277against the llama.cpp BF16 GGUF reference020040060017212564149155164292292595Q8_0vLLM vs llama.cppQ6_KQ5_K_MIQ4_XSQ4_K_SQ4_K_MQ3_K_MIQ3_MQ2_K

Top: vLLM output throughput as concurrency rises on a mixed request shape. The dashed line is this page's BF16 GGUF on the A100 at a single stream, 251.1 tokens per second, drawn at one request. Bottom: decisions changed out of 1,277 confirmatory prompts, each counted against the llama.cpp BF16 GGUF.

ConcurrentShapeRequests/sOutput tok/sTotal tok/s
1long8.2139.142,977
1median5.3270.54,455
1mixed8.5279.6407
1short6.8225.2328
8mixed41.71158.51,827
32mixed118.23323.55,901
64mixed208.45814.612,043
128mixed269.67794.341,524
256mixed259.98312.775,138

Requests per second peak at 128 concurrent and dip slightly at 256, while total tokens per second keeps climbing — the later gain is prompt tokens, not decode. One request of each shape was also run separately: long 139, median 271, mixed 280, short 225 output tokens per second. Cold start, loading the checkpoint into vLLM, was 52.6 seconds.

What the engine changed

Changing the engine, the GPU and the kernels together moved 21 of 1,277 decisions — 98% agreement, and 87% on exact output. Of those flips, 3 fall on the six prompts already known to tokenize differently between the two engines, leaving 18 attributable to engine numerics. For scale, the Q6_K quantization this page recommends changed 25. Moving between inference engines cost about as much per-example disagreement as the quantization OpenGrad spent a phase gating, which is why the comparison is reported here rather than folded into the ladder. The vLLM rerun reproduced the frozen reference within 0.011 call F1 at a parse-valid rate of 100%, so both engines were scored against a confirmed reference. Method and the per-example flips are in the H200 run report and the agreement record.

Conditions

Enginellama.cpp d3146f2, CUDA backend, llama-bench
HardwareOne NVIDIA A100-SXM4-80GB, on Modal
Offloadall layers on the GPU
Batching2048 batch, 512 micro-batch, 24 CPU threads
KV cachef16 keys, f16 values
Repetitions3 per test
Run date2026-09-12
Engine, vLLM runvLLM 0.29.0, torch 2.13.0+cu130, CUDA 13.0
Hardware, vLLM runOne NVIDIA H200, 139.8 GiB, on Modal
Batching, vLLM run5760-token context, up to 256 sequences, prefix caching off, bfloat16
Run date, vLLM run2026-09-12

The throughput rows are single-stream llama-bench runs from an empty context. The vLLM run is a separate batching sweep on different hardware, so the two are not comparable row for row. Laptop, phone and CPU throughput were not measured for these files; on a phone, where memory bandwidth is the limit, the ranking can differ.