Reference
Benchmarks and evals
Measure variants against each other honestly, score models on prompt suites, and put a perplexity number on held-out text.
Three numbers in Kvad's own documentation were once wrong, each from timing a single run. That is why the server has a benchmark in it rather than a shell script, and why it is strict.
#Benchmarks
kvad bench run Qwen/Qwen2.5-1.5B-Instruct@cpu-q8 Qwen/Qwen2.5-1.5B-Instruct@gpu-q8
kvad bench # runs
kvad bench show 4| Option | |
|---|---|
--rounds N |
How many rounds. |
--tokens N |
Tokens to generate each time. |
--prompt T |
The prompt. |
--seed N |
What a run does, and why:
- A round visits every variant once, and a run is several rounds. Five of A and then five of B blames the model for whatever changed about the machine in between.
- Median and range, from samples that are all kept. Two ranges that overlap are two numbers that have not been told apart.
- The KV cache is dropped before every timed generation, or the second round reports a time to first token that no first run would ever see.
- It will not start while a training run, an eval or another benchmark is going. The refusal names what is in the way.
One limit: interleaving is right when the variants are small enough to be in memory together. Two variants that are each a third of the machine's memory evict each other, and each round then measures the other's footprint. Run those as two runs.
The web UI's Benchmarks page starts the same runs and draws every sample.
#Prompt suites
A suite is a file of prompts, each with what a right answer contains:
{
"name": "capitals",
"cases": [
{ "prompt": "What is the capital of Norway?", "expect": "Oslo" },
{ "prompt": "What is the capital of France?", "expect": "Paris" }
]
}kvad evals add capitals.json
kvad evals suites
kvad evals run capitals Qwen/Qwen2.5-1.5B-Instruct Qwen/Qwen3-14B@gpu-q8
kvad evals show 2Each model is loaded in turn and every verdict is kept. A model with no
@BACKEND runs on --backend, or on the server's default.
#Perplexity
Perplexity scores how well a model predicts text it has not seen. Lower is better. It is the number to use when comparing two training runs on the same corpus, or the same model at two precisions.
kvad datasets add heldout.txt
kvad evals perplexity heldout my-model my-model-v2Scoring runs the output head over a whole window as one matrix product, so it is about four times as fast as decoding the same number of tokens.