Use

The terminal app

Browse, download, switch backend and chat, without leaving the terminal.

kvad-tui

kvad-tui is on macOS only. It loads models in its own process, so it needs no server; if the background service is holding a large model, unload it first or the two will compete for memory.

#Keys

Key
/ Search Hugging Face.
↑ ↓ Select a model.
enter Download it if needed, and load it.
p Cycle the backend.
d Delete the selected model.
tab Switch between the model list and the chat.
esc Interrupt a generation.

#Trying the backends

p cycles all six backends: cpu f32, cpu q8, cpu q4, gpu bf16, gpu q8, gpu q4. It is the quickest way to feel the trade-offs. Load a model at f32 and ask it something arithmetic, reload at q4 and ask again, then reload on the GPU and watch the token rate.

#The status bar

After each reply the status bar reports the decode rate and how much of the conversation came from the cache:

[33.3 tok/s · 96 cached]

The chat keeps its KV cache across turns. Each turn the conversation is re-encoded, compared with what is cached, and only the new message is prefilled. Without that, turn N would re-read the whole transcript.

esc

Try , or .