Use
Chat and completion
One-shot answers and conversations from the command line, the sampling options, and what the numbers after a reply mean.
#One answer
kvad run --prompt "Explain backpropagation in one sentence."llama · 30 layers · 9 heads (3 KV heads, 3x grouped) · 576 embd · 8192 ctx
134.5M parameters · instruction-tuned
The sky appears blue because of the way our eyes detect light. [...]
[prefill 41 tokens in 1.10s · generated 56 in 1.62s = 34.6 tok/s]An instruction-tuned model gets its own chat template around the prompt. A
base model continues the text. To continue the text with an instruct model
too, on a server, add --raw.
#A conversation
kvad chat
kvad chat --system "You are terse."
kvad chat --model Qwen/Qwen3-14BOn a server, --save keeps the conversation there, where the web UI shows it,
and --conversation ID carries on with a kept one:
kvad chat --save
kvad conversations # ls
kvad chat --conversation 12#Sampling
| Option | Default | |
|---|---|---|
--max-tokens N |
256 | The generation budget. |
--temperature F |
0.7 | 0 is greedy. |
--top-k N |
40 | Keep the N best candidates. |
--top-p F |
0.95 | The nucleus threshold. |
--seed N |
7 | The same seed draws the same reply. |
--greedy |
Shorthand for --temperature 0. |
#The numbers after a reply
Prefill is reading the prompt: every token of it goes through the model once, in batches. It is bound by arithmetic, which is what a GPU is for.
Decode is writing the answer, one token at a time. It is bound by how fast the weights can be read from memory, which is why fewer bits per weight helps and why the CPU at q8 can keep up with the GPU at bf16.
Cached tokens are the part of the conversation the model had already read. A conversation keeps its KV cache across turns: each turn the new transcript is compared with what is cached, and only the new message is prefilled.
#Reasoning models
Qwen3 and its relatives think before they answer, between <think> and
</think>. Kvad separates the two as they stream. The command line prints the
working to stderr and the answer to stdout, the web UI shows the working
collapsed above the answer, and the API sends it as reasoning_content,
leaving content the answer.
A small reasoning model can spend its whole budget thinking. If a reply is
empty, raise --max-tokens.
#Where it runs
With the service running, kvad run and kvad chat send their work to it, so
a second copy of the weights is not loaded beside the one the server holds.
The first line of output says which happened.
kvad chat --local # in this process, whatever is running
kvad chat --remote http://gpu-box.local:5823See The server for the order these are decided in.