Serve
The server
One binary that serves the API and the web UI, holds the models, and runs everything long as a job.
kvad serve # loopback on 5823, no authentication
kvad serve --bind 0.0.0.0:5823 # needs an auth mode, or --insecurekvad serve hands over to kvad-serve, which is one file with the web UI
built into it. Run it by hand like this, or let the
background service start it at login.
| Option | |
|---|---|
--bind ADDR |
The address to listen on, such as 127.0.0.1:5823. |
--config PATH |
The configuration file. Default ~/.config/kvad/kvad.toml. |
--db PATH |
The SQLite database file. |
--insecure |
Allow a non-loopback address with no authentication. |
#What it holds
- Models in memory. Several at once when they fit, each charged for its weights and a full context of KV cache. See what is in memory.
- A queue. One scheduler owns the engine, and every generation waits its turn behind it. The queue depth is on the dashboard.
- Jobs. Training runs, evals, benchmarks, crawls and videos are jobs: they answer at once, carry on if the client goes away, and can be followed or cancelled.
- A database. SQLite, for what the filesystem cannot answer: conversations, the settings of every image and video, training runs, accounts and API keys.
#One generation at a time
The engine runs one generation at a time. Requests queue, and the server is built to say so rather than hide it:
- The playground's side-by-side comparison and the benchmark page both load each variant in turn, and say so on the page.
- A benchmark will not start while a training run, an eval or another benchmark is going, and the refusal names what is in the way.
- Chat during a training run is slower, not blocked: roughly half speed, measured.
#Loading at startup
By default the server holds nothing until something asks, so the first request after a reboot pays for the load: tens of seconds for most models. To have the default model loaded as soon as the server is listening:
[server]
autoload = trueIt is off by default because the cost is a model in memory a minute after boot whether or not anybody turns up.
#The command line is a client
With a server running, kvad sends its commands to it, so the command line
and the server share one copy of the weights. The first line of output says
where a command went:
$ kvad chat
kvad: http://127.0.0.1:5823 (the kvad service; --local to run in this process instead)
model: Qwen/Qwen3-14B@gpu-q8A command is sent to the first of these that is set:
--remote URL, or--localto run it in the process.KVAD_URLin the environment.urlunder[client]inkvad.toml.- This machine's service, if it answers.
The first three are things you said, so a server you name that does not answer is an error. The last is a guess: when nothing answers there, the command runs in the process.
That makes a Mac in another room a server for a laptop:
export KVAD_URL=http://mac-studio.local:5823
kvad auth login
kvad chat#Health
kvad service status
curl http://127.0.0.1:5823/api/health/api/health says what is loaded, how deep the queue is, and which
authentication mode is in force. kvad metrics prints the machine's memory,
the KV cache, and recent request timings.