Apple silicon

Built on a MacBook Pro, for the chip inside it.

Kvad is developed and measured on an M5 Pro. Its GPU engine is written for Metal and nothing else, and it has kernels of its own for the part of the M5's GPU that general-purpose shaders never touch.

The matrix units

The M5's GPU carries hardware for matrix products. An ordinary Metal compute shader does not use it: it is reached only through Metal 4's tensor operations. The framework Kvad's GPU engine is built on stops at about 7.5 TFLOP/s on this machine whatever it is given, because its shaders are the ordinary kind.

So Kvad has its own. One kernel takes dense half-precision weights. Another takes quantised weights, unpacks each slab on the GPU, and hands it to the same hardware, which matters because the models worth running locally are the quantised ones.

They are compiled when a model loads and used when the GPU reports the Apple10 family. On an M1 to M4 the same model takes the standard path, with nothing to configure.

Per matrix productAgainst the stock kernel
Dense, 512 rows or more3.0–3.9×
Dense, 16 to 128 rows2.3–3.9×
Dense, 8 rows1.2–1.6×
Quantised, 8-bit2.3–2.7×
GGUF k-quants (Q4_K, Q5_K, Q6_K)75–90% of the 8-bit rate
One row, as in decode0.87–0.95×, so not used

Timed in alternating rounds against the stock kernel, because this machine drifts by up to a third between runs. The probe and its results

What that is worth end to end

A model is more than its matrix products, so the whole-step numbers are smaller than the kernel's. These are the ones that matter.

On an M5 ProBeforeAfter
Prefill, Qwen2.5-1.5B at bf16, 2,079 tokens0.99 s0.39 s2.5×
FLUX.1-schnell at q8, a step at 1024²12.9–13.7 s8.7–9.2 s1.5×
SDXL at f16, a step at 1024²3.90–3.98 s3.09–3.15 s1.26×
Qwen-Image with a LoRA applied, a step4.5% slower than without one

SDXL gains least because most of its time is convolution, which these kernels do not touch.

What did not get faster

Decoding. Writing a reply is one token at a time, and each token reads every weight in the model once. That is bound by how fast memory can be read, not by arithmetic, and no matrix unit changes it. On one row the new kernel is slower than the stock one, so Kvad does not use it there.

What helps decode is fewer bytes per weight and fewer trips to the GPU. Quantising to 4 bits does the first. Fusing a layer's small operations into three Metal kernels does the second, and took a 1.5 B model's layer from 25 kernel launches to 10. At q8 that is 114–117 to 126–129 tokens a second early in a conversation, and 102–108 to 110–115 with 2,048 tokens of context.

The benchmark on the right is from the day this page was written. The GPU at q4 decodes 4.4 times as fast as the CPU at q8, and 1.6 times as fast as the GPU at q8, for the same model.

$ kvad bench run Qwen/Qwen2.5-1.5B-Instruct@cpu-q8 \
      Qwen/Qwen2.5-1.5B-Instruct@gpu-q8 Qwen/Qwen2.5-1.5B-Instruct@gpu-q4
job 2 — 3 variants, 5 rounds (bench)

MODEL                                RUNS  DECODE TOK/S, MEDIAN (RANGE)  FIRST TOKEN MS
Qwen/Qwen2.5-1.5B-Instruct · cpu-q8  5     48.5 (47.2–49.3)              72.3
Qwen/Qwen2.5-1.5B-Instruct · gpu-q8  5     132.5 (131.1–133.9)           34.8
Qwen/Qwen2.5-1.5B-Instruct · gpu-q4  5     215.4 (213.2–218.3)           34.1

Unified memory is the budget

On a Mac the GPU and the CPU share one pool, so the question is never how much VRAM a card has. It is how much of the machine a model may take. By default Kvad lets models use three quarters of it and leaves the rest to you.

A model is charged for its weights and for a full context of KV cache before it is admitted, so one that loads can always finish. One that does not fit is refused, by name, rather than loaded into swap.

The table is what each model holds once loaded, measured. Add the KV cache for a language model: Qwen3-14B is charged 5 GB for 32,768 tokens of it.

ModelPrecisionIn memory
Qwen2.5-0.5Bq80.6 GB
Qwen2.5-1.5Bq81.6 GB
Stable Diffusion 1.5f161.9 GB
SDXLf166.4 GB
Qwen3-14Bq814.6 GB
DeepSeek-V2-Lite, 15.7 B parametersq817.7 GB
FLUX.1-schnellq818.3 GB
Qwen-Image, 20 B parametersq829.2 GB

The CPU is not an afterthought

Half of Kvad is a CPU engine with no framework under it: block-wise 8-bit and 4-bit weights, an integer dot product on Arm's i8mm instructions, a hand-written 4-bit decode kernel, and a tiled float matrix product.

On an M5 Pro it decodes Qwen2.5-0.5B at 112 tokens a second, and DeepSeek-V2-Lite, 15.7 billion parameters, at 23. It is also what a Linux machine gets.

How the CPU engine works
Decode on the CPU, q8Tokens a second
SmolLM2-135M194
GPT-2 medium128
Qwen2.5-0.5B112–117
Qwen2.5-1.5B48.5
DeepSeek-V2-Lite, 15.7 B parameters23–27

An M5 Pro, 18 cores. The ranges are separate measurements on different days.

Which Mac

ChipWhat you get
M5 familyEverything on this page. This is what Kvad is measured on.
M1 to M4The same models and features through the standard Metal kernels. Slower prefill and image steps; decode is much the same.
IntelNo build.

Every number here is from one machine, an M5 Pro with 48 GB in a MacBook Pro. Kvad has not been measured on an M5 Max, a Mac Studio or a Mac mini, and this page will say so when it has. If you run it on one, the benchmark page will give you numbers worth sending in.

Try it on yours.

Platform details · Run the benchmark yourself

esc

Try , or .