How it works
kvad-gpu — the same model, handed to a framework
kvad-gpu run --prompt "Why is the sky blue?"
kvad-gpu run --device metal --dtype bf16 --model Qwen/Qwen2.5-0.5B-Instruct
kvad-gpu chatmodel.rs is the Llama forward pass again, in
candle. Read it next to
llama.rs: the structure is line for line the
same, and every operation is one you already wrote.
It is three forward passes now rather than one.
gpt2.rs and
deepseek.rs sit beside it, with
common.rs holding what they share — which is a
shorter list than it looks: two weight layouts behind one forward, an
embedding table that is dense or quantised, and the reader every checkpoint is
loaded through. session picks which one from the config, and that same list
is what the server asks before it offers the GPU as a default, rather than
keeping a copy of the answer.
matvec_btbecomesTensor::matmul.- The loop over attention heads becomes one batched matmul over a
[1, n_head, seq, head_dim]tensor. candle_nn::rotary_emb::ropeis the samerotate_halfconvention asRope::apply, andops::rms_normthe same scale-only normalisation.
Two differences are worth noticing because they run against the framework:
- Batching is free. The CPU engine needed a separate
forward_batch, because a matrix-vector product and a matrix-matrix product are genuinely different kernels. On the GPU there is oneforward, and a single token is just the narrow case. - The causal mask comes back. With a KV cache the CPU engine got masking
for free — the cache only ever held earlier positions. Here the whole batch
is one matmul, so the future has to be masked out explicitly, with the
triangular
-infmatrix every tutorial shows.
#Numbers
| CPU (q8) | Metal (bf16) | ||
|---|---|---|---|
| Qwen2.5-0.5B, decode | 112 tok/s | 113 tok/s | 1.0x |
| SmolLM2-135M, decode | 194 tok/s | ~168 tok/s | 0.9x |
| Qwen2.5-0.5B, prefill (640 tokens) | 1.16 s | 0.19 s | 6.1x |
The GPU wins on prefill and, at this model size, not at all on decode. That is the same distinction as everywhere else in this repo. Prefill is compute-bound, which is what a GPU is for. Decode is memory-bound, and there the CPU is carrying int8 weights (556 MB) against the GPU's bf16 (1260 MB) — quantisation claws back a whole hardware generation, and on the smaller model the CPU is ahead.
That row used to read 35 tok/s and 3.3x, and the smaller model used to be slower on the CPU than the larger one. Both were a scheduling bug, not a property of the hardware. It is the same lesson as the rest of the repo, learned the expensive way: the number you are explaining may not be a number about the thing you think it is about.
Metal f32 runs at about two thirds of bf16's rate, for exactly double the
memory. bf16 is the default because it is also what the checkpoints ship as.
The GPU backend covers the Llama family, GPT-2, DeepSeek V2, Qwen3.5/3.8 and
Qwen3-Next — five of the six architectures the CPU engine has. deepseek_v3
is the sixth: deepseek.rs implements it and
session has no arm that routes a V3 checkpoint to
it, so asking for one on the GPU says which architectures there are rather
than failing obscurely. One list answers that, and the server asks it rather
than keeping a copy.
#Quantised weights on the GPU
kvad-gpu run --quant q8 # or q4, q4k, q6kcandle's QTensor holds GGML's block formats — the same scheme as
quant.rs: blocks with a shared scale, Q8_0 and
Q4_0 being 32-wide exactly like ours. QMatMul keeps HuggingFace's
[out, in] layout and transposes inside its kernel, where the dense path wants
the transpose done once at load, so Proj hides the
difference.
All six backends, Qwen2.5-0.5B, decode:
| backend | weights | tok/s |
|---|---|---|
| cpu f32 | 1976 MB | 64 |
| cpu q8 | 556 MB | 112 |
| cpu q4 | 309 MB | 116 |
| metal bf16 | 1260 MB | 113 |
| metal q8 | 525 MB | 191 |
| metal q4 | 278 MB | 228 |
Quantisation is worth 1.7x on both now — 113 → 191 on the GPU, 64 → 112 on the CPU. It used to be worth nothing on the CPU, and the explanation given here (a per-matmul cost that the bytes could not touch) was correct about the mechanism and wrong about the conclusion: that cost was removable, not inherent. With it gone, both backends are bandwidth-bound in decode and behave the same way.
And then prefill, where it stops being free — eventually:
| prompt | metal bf16 | metal q8 | |
|---|---|---|---|
| 140 tokens | 0.070 s | 0.040 s | q8 1.75x faster |
| 640 tokens | 0.180 s | 0.160 s | q8 1.12x faster |
| 1340 tokens | 0.390 s | 0.380 s | even |
| 5040 tokens | 2.590 s | 2.750 s | q8 1.06x slower |
There is a crossover, at around 1300 tokens on this machine. A short prompt is still narrow enough that reading the weights dominates, so smaller weights win; a long one turns every matmul into real work over a wide batch, and then the dequantisation is pure overhead. The bottleneck moves with the batch size, and the point where it crosses is a property of this GPU rather than of quantisation.
(The CPU engine once claimed a similar crossover against matrix size, and that one turned out not to exist — it was measuring a scheduler. This one is measured across four prompt lengths with five runs each, which is the standard the other claim failed.)
Correction. An earlier version of this section claimed a flat "1.6x slower on prefill" from a single pair of measurements at 654 tokens (0.24 s against 0.39 s). That does not reproduce: five runs at each of four prompt lengths give the table above, with a median equal to the minimum every time. Re-running the old arrangement behind
KVAD_GPU_DENSE_EMBED=1rules out the embedding change as the cause, so the original figure was simply a bad measurement on a loaded machine. The shape of the claim survives — there is a length past which quantising costs you — but it arrives much later and much more gently than reported. Measure the curve, not two points on it.
One practical note. The k-quants (q4k, q6k) need dimensions divisible by
256 — Qwen is 896 wide, so they are refused up front with a message naming
q8 and q4 instead of failing halfway through the load.
#The table that was stored twice
kvad-gpu run --quant q8 # quantised table
KVAD_GPU_DENSE_EMBED=1 kvad-gpu run --quant q8 # the dense one, for comparisonThe embedding table was the last dense tensor in a quantised GPU model, and on
a small model it is not a small one: 136M of Qwen2.5-0.5B's 494M parameters live
in [151936, 896]. At bf16 that is 272 MB.
It stayed dense because the lookup needs an operation nothing else in the engine
wants — gather rows and dequantise only those. A QTensor is a run of blocks,
not a matrix, so row t is a span of blocks that has to be decoded on its own.
candle turns out to ship exactly that kernel (QTensor::embedding, which is
GGML's get_rows), so the change is one method call rather than a Metal shader.
The compression is the smaller half of it. Qwen ties its embeddings — small
models nearly always do — so the lookup table and the output head are the same
matrix. They were stored twice anyway, because a dense index_select wants
[vocab, n_embd] and a matmul wants the transpose. Quantise the lookup and both
want the identical thing: one Arc<QTensor>, two uses.
| before | after | |
|---|---|---|
| metal q8 | 797 MB | 525 MB |
| metal q4 | 550 MB | 278 MB |
525 MB for 494M parameters is 8.5 bits each, which is exactly Q8_0 — 8-bit
codes plus an f16 scale per 32. The model is now at the format's floor, with
nothing left dense but the norms, and the win over bf16 goes from 1.6x to
2.4x (4.5x at q4).
And it changed the speed by nothing at all: 196.6 → 195.7 tok/s at q8, 239.6 → 237.3 at q4, prefill identical to three decimals. Exactly as it should be. One row of 151936 is read per token, so the table was never on the bandwidth path — it was pure capacity. Everywhere else in this repo, quantisation traded accuracy for speed. Here it trades accuracy for room, and that is the whole point: memory is what stops you loading a bigger model, and a bigger model is worth far more than 1% of tok/s.
The CPU engine needed no change for any of this. Weight::row has dequantised
one row at a time since quantisation was added, and the tied head has always
been the same Weight — the hand-written version got here by the path of least
resistance, because one type did both jobs. The framework version duplicated
precisely because its two operations wanted two different types.
#The quantising, done once over there too
The CPU engine got a cache of pre-quantised weights some way back, and this backend did not, because candle quantises Qwen2.5-0.5B in about 0.2 s and 0.2 s is not a problem worth a file format. That judgement was correct and it did not scale: the same code on DeepSeek-V2-Lite spends 66 seconds turning 15.7 billion bf16 weights into 4-bit blocks, every single load, and throws the result away on exit.
So crates/gpu/src/qcache.rs writes them out.
Steady state, three runs each way, both orders:
| model | quantise | map | file | |
|---|---|---|---|---|
| GPT-2-medium q8 | 1.7 s | 0.8 s | 358 MB | 2.1x |
| Qwen2.5-0.5B q8 | 1.2 s | 0.9 s | 500 MB | 1.3x |
| Qwen3-0.6B q8 | 1.3 s | 0.9 s | 604 MB | 1.4x |
| DeepSeek-coder-7B q8 | 6.6 s | 2.8 s | 6.8 GB | 2.4x |
| DeepSeek-V2-Lite q4 | 65.5 s | 3.9 s | 8.2 GB | 17x |
Those are whole-process times and the small ones are mostly floor — about 0.7 s of it is the Hub client resolving five file paths, which the CPU section already noticed is the thing left standing once the weights stop costing anything. The row that matters is the last one.
Same box, different contents. The file is the same .nq container: magic,
data section, a JSON header at the end, a u64 saying where the header starts.
What goes in it is not the same bytes. Both engines call their eight-bit format
q8 and they disagree about what that means — ours carries an f32 scale per
32 weights in the order quant.rs reads, GGML's carries an f16 — so handing
one to the other would be a model made of noise. They are told apart by name
before anything else: Qwen--Qwen3-0.6B.q8.nq and
Qwen--Qwen3-0.6B.gpu-q8.nq, in one directory, both listed by kvad cache and
both forgotten by kvad cache <repo>. Container and Writer moved out of
qcache.rs to be shared, which is the whole of the code reuse: a second copy of
the trailer arithmetic is a second place for it to be wrong.
The part that got simpler. The CPU cache has to stamp the checkpoint's
entire tensor list into every file, because a mapped cache never reopens a
checkpoint — so a build that learns to read one more optional weight gets
None back and runs without it, at full speed, saying nothing. That was
the last route by which an unread weight could hide.
This one has no such problem, and the reason is that it caches less. Only the
quantised matrices go in the file. The norms, the biases and DeepSeek's two
dense halves of kv_b_proj still come from the checkpoint, which therefore is
still open, which means a name the file does not hold is simply quantised from
the source on the spot. A miss is answered rather than survived. The file is
then deleted at the end of that load, so the next one writes a complete one, and
the unread-tensor guard goes on subtracting from the checkpoint's list exactly
as it did before — a cache hit records the name it answered, so a weight served
from disk still counts as read.
That also means there is nothing to cache at --quant none: the checkpoint
already is the weights, and a copy of them would be a copy of the checkpoint.
The same rule the CPU cache applies at f32.
Where the win is, and is not. It is in the arithmetic, not the I/O. Run the two arrangements alternately on a machine whose page cache cannot hold both the 29 GB checkpoint and the 8.2 GB cache file, and V2-Lite still wins (66 s to 17 s) while DeepSeek-coder-7B comes out even: the smaller file is being read cold either way, and 6.9B parameters is only a few seconds of quantising to save. Load the same model twice in a row — which is what a cache is for — and the table above is what you get.
#One trait, two backends
Adding the GPU needed a change to the CPU engine's shape.
Transformer deliberately keeps the KV cache
outside the model, because on the CPU it is a Vec<f32> the caller can own.
A GPU backend cannot work that way — its cache lives in device memory and must
never round-trip through the host between layers.
So the cache moved inside, behind a Session trait: forward, cached,
truncate. Everything above it — prefix reuse, sampling, streaming, the chat
loop, the TUI — stopped caring which backend it was talking to. The CLI keeps
its --quant, the GPU CLI gets --device and --dtype, and in the TUI p
cycles through all six: cpu f32, cpu q8, cpu q4, gpu bf16, gpu q8,
gpu q4.
That is the usual shape of this kind of work: the interesting part was not the new backend, it was the seam the old one had to grow.
#The seam that cannot rewind
truncate returns a number now, and the number is the interesting part.
Every architecture released since Qwen3 — Qwen3.8, Qwen3-Coder-Next,
GLM-5.3-Flash, Kimi-Linear — carries a recurrent state on three layers in
four, with ordinary attention on the fourth. A state is not a cache. Keys and
values are stored per position, so the state of a prefix is exactly the rows
belonging to that prefix and dropping the rest is a Vec::truncate. A
recurrent state is one fixed-size vector that has already absorbed every token
it has seen, and there is no subtraction that takes the unwanted ones back
out.
Prefix reuse is built on being able to take them back out. So this engine's answer is to refuse: a cache with a recurrent state rewinds to zero or not at all. Every turn of such a conversation is a fresh prefill — honest and slow, rather than fast and wrong.
Which is why truncate reports where it actually landed instead of returning
(). The caller used to ask for a rewind to N, assume it happened, and forward
from N. On a state that refused, that would run the new tail over a state that
had already read it: a model that is wrong rather than slow, and wrong in a way
nothing prints. Now the caller forwards from the position it is given and
reports that as cached_tokens, so the refusal shows up as a number on the
dashboard rather than as a model quietly talking nonsense.
The alternative is to snapshot the state every N tokens and rewind to the nearest one, at a memory cost proportional to context / N. That is a real design with a real tuning knob, and nobody can choose N without a model to measure. Returning a number is what makes it a change to one cache later instead of a change to everything that calls it.
Counting what the cache holds turned up a sevenfold error in what the server
was reporting. kv_bytes_per_token read kv_dim() — what ordinary
attention would store — where it should have read what the cache actually
stores. Those agree for GPT-2 and Llama and disagree for DeepSeek, which sets
n_kv_head = n_head because every head really does have its own key, and then
keeps none of them: 576 floats a position against the 2048 kv_dim()
describes. The server was reporting 7.1x the memory it was using, on the
one architecture whose entire argument is that it uses less.
#Pictures, which are not tokens
cargo run --release -p kvad-gpu --example sdxl -- \
--prompt "a lighthouse on a cliff at dusk, oil painting" --out lighthouse.png
kvad images make "a red fox in fresh snow" --model stabilityai/stable-diffusion-xl-base-1.0image/ is text-to-image, and it is the one
part of the engine with no CPU version. That is a decision, written down with
every shape it depends on in docs/image-plan.md: an
image model is a convolution stack, tensor.rs has no convolution, and a fast
one is a project of its own. What is still written here is the model. CLIP,
the UNet, the VAE, both schedulers, Qwen-Image's MMDiT and its video VAE,
FLUX's T5 and single-stream blocks, and the PNG encoder are all in this
repository; candle supplies conv2d and the
matmuls, and nothing from candle-transformers is used.
Nothing about it is a token loop. A text encoder runs once. A denoiser runs
20–50 times from noise, twice per step while guidance is on, and a scheduler
with no weights at all turns each prediction into the next latent. A VAE turns
the last latent into pixels. So a painter is not a Session: it implements
kvad::image::Painter, a peer of Llm that the
engine thread, the scheduler and the server hold beside it.
The first image out of examples/sdxl.rs was the right one: a lighthouse,
1024², 30 steps. What that took on an M5 Pro, in f16, from single runs (the
second with a 58 GB download going on in the background):
| 1024², 30 steps | 768², 20 steps | |
|---|---|---|
| denoise, per step (guidance on) | 3.93 s | 2.2 s |
| VAE decode | 9.4 s | 5.3 s |
| load, from the disk | 19 s | — |
The decode is the most memory the pipeline ever asks for: it is the only stage at full resolution, 128 channels of 1024×1024.
Qwen-Image is the same story at twenty billion parameters: Qwen2.5-VL-7B as the text encoder, a 60-block MMDiT with a three-axis RoPE as the denoiser, a video VAE run on one frame, flow matching instead of noise prediction. 41 GB of bf16 denoiser does not fit beside anything on a 48 GB machine, so it runs at q8 through the same loader and quantised-weight cache as the text models: 21.7 GB of transformer and 7.5 GB of language tower, quantised once (130 s with the quantising, 66 s from the cache after). Again single runs, true CFG at 4:
| 512², 20 steps | 1024², 20 steps | |
|---|---|---|
| denoise, per step (two 20B passes) | 6.5 s | 28.3 s |
| VAE decode | 1.5 s | 5.8 s |
| peak memory footprint | 32.7 GB, quantising as it loaded | — |
Its first image was structured noise, and the model was not the reason. The
language tower checked out at once — it is a chat model, so the checkpoint's
own output head could be pointed at it, and "The capital of France is" came
back " Paris". The denoiser was probed instead: from pure noise at σ = 1, the
clean latent it predicts should be smooth, and its neighbouring pixels
correlated at 0.65. candle's Metal kernel for a quantised matrix-matrix product
reads its input from the start of the buffer whatever the tensor's offset — and
the image half of joint attention is a narrow that starts where the text
ends. So every patch was projected from the wrong rows. Tensor::copy keeps the
offset too; force_contiguous does not. With that in Proj::forward the same
probe read 0.996, and the next image was a fox.
FLUX.1-schnell came after, and needed less new code than either: its first
nineteen blocks are Qwen-Image's blocks under other names, so the block moved
into mmdit.rs and both models load it. What
FLUX adds is a T5-XXL encoder (t5.rs: relative
position buckets instead of positions, no 1/√d on the scores) and 38
single-stream blocks that run attention and MLP side by side. Its first image
was right. At q8 it holds 18.3 GB (12.6 GB of transformer, 5.1 GB of T5), and
schnell's four steps need no guidance, so each is one forward pass:
| 768², 4 steps | 1024², 4 steps | |
|---|---|---|
| denoise, per step | 7.3 s | 13.1 s |
| VAE decode | 5.3 s | 9.6 s |
(Single runs again. FLUX.1-dev takes guidance as an input and is refused by name until it has an implementation of its own.)
The community's GGUFs of Qwen-Image's and FLUX.1-schnell's transformers load
too, and of LTX-2.5's distilled DiT for its fast pipeline, by the repo and
the quantisation together; the model card's base_model supplies the rest
(docs/gguf-plan.md):
kvad pull city96/Qwen-Image-gguf:Q4_K_S
kvad images make "a red fox in fresh snow" --model city96/Qwen-Image-gguf:Q4_K_Scity96's FLUX files keep Black Forest Labs' layout, one qkv matrix where
diffusers has three, and are read through a map of rows, so no block is
requantised. city96's Q8_0 of either model is Kvad's own q8 block for block.
The k-quants (Q4_K, Q5_K, Q6_K) run on the M5's matrix units in the same
kernel as Kvad's q8, one decoder a format, at 75–90% of its rate. So a
Q4_K_S holds 9.1 GB less for Qwen-Image and 5.5 GB less for FLUX at 1024²
for a step 5–18% longer than q8's, and an LTX-2.5 Q4_K_M holds about 4 GB
less and denoises as fast.
SDXL fine-tunes load too: the diffusers folders nearly all of them ship,
whichever names their files carry, and a checkpoint in one file in
Stability's own layout, as Pony and Civitai's are. A repo whose only model
is one file is named by the repo; one with several, repo:file.safetensors;
a file on this machine, by its path. Its configs come from SDXL's base, and
its VAE is the fp16-fix (docs/checkpoint-plan.md):
kvad pull LyliaEngine/Pony_Diffusion_V6_XL
kvad images make "score_9, a red fox in fresh snow" --model LyliaEngine/Pony_Diffusion_V6_XL
kvad images make "a lighthouse at dusk" --model ~/Downloads/some-sdxl-finetune.safetensorsStable Diffusion 1.5 is SDXL's parts at a smaller size: one text encoder, read to its last layer; a UNet of four levels with no size conditioning, whose transformers project with 1×1 convolutions; and its own VAE, which, unlike SDXL's, stays in range in f16. Each part matches diffusers at 112–121 dB in f32 (docs/sd15-plan.md). It is here for its fine-tunes, which load by the same three kinds of name as SDXL's, and from one file its VAE is the file's own, since a fine-tune's often is: all 248 of DreamShaper 8's VAE tensors differ from the base's. At 512² and 25 steps it takes 0.75 s a step on an M5 Pro, and the server charges 1.9 GB.
kvad images make "a red fox in fresh snow" --model stable-diffusion-v1-5/stable-diffusion-v1-5
kvad pull ckpt/anything-v5.0
kvad images make "1girl, reading in a library" --model ckpt/anything-v5.0LoRAs apply to Qwen-Image, FLUX, SDXL, SD 1.5 and LTX-2.5, at run time, per
request: the model stays as it was loaded, q8 or a GGUF, and each request
chooses its own LoRAs and strengths (docs/lora-plan.md).
Each is a side path (x·A)·B beside a layer it adapts, added into the
layer's answer by the M5 dense kernel's store, for 4.5% a step on Qwen-Image.
They are read in the spellings people publish: PEFT's, kohya's under ldm's
names, diffusers' or Black Forest Labs' (whose fused qkv is split by rows),
diffusers' two older ones, with their text encoders' and their convolutions'
pairs. Against diffusers with PEFT, Qwen-Image's first two blocks with
Lightning agree to 95.8 dB in f32, three FLUX LoRAs to 111 dB, and seven SDXL
and SD 1.5 LoRAs of every one of those kinds to 108–121 dB. On LTX-2.5 a
video's LoRAs go on every DiT its pipeline loads, and in bf16 they are level
with the reference's own run of its distilled LoRA. Lightning's 8 steps take
68 s at 1024², where the base's 50 took 448 s, and draw a sharper picture:
kvad pull lightx2v/Qwen-Image-Lightning:Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors
kvad images make "a tiny astronaut hatching from an egg on the moon" --model Qwen/Qwen-Image \
--steps 8 --lora lightx2v/Qwen-Image-Lightning:Qwen-Image-Lightning-8steps-V2.0-bf16.safetensorsA LoRA for SDXL can be trained here too, on a folder of pictures with a caption beside each (docs/tune.md): 1.9 s a step in 7 GB at 512², 7.3 s in 14 GB at 1024², keeping the step with the lowest loss on pictures it did not train on.
kvad-gpu tune --data ./my-photos --name my-styleTwo things in it are measurements rather than code:
- The previews. Each step can carry a picture of where it is heading, without a decode: the latent's four channels mixed into RGB by a fixed 4×3 matrix. The matrix was fitted by least squares from one generated image, latent against pixels, and explains 76–83% of the variance per colour — blurry and slightly wrong, which is what a preview is for.
- The schedule. SDXL's Euler schedule is checked against the timesteps
diffusers picks (
958, 925, …, 34, 1for 30 steps), because a scheduler that visits the wrong noise levels still draws a picture, only a worse one.