Use
Images
Text to image with Stable Diffusion 1.5, SDXL, FLUX.1-schnell and Qwen-Image, on the Mac's GPU.
Image generation runs on the GPU engine, so it needs a Mac. The models are written out in this repository: CLIP, T5, the UNet, the MMDiT, the VAEs and both schedulers. Only convolution and matrix multiplication come from a framework.
kvad pull stabilityai/stable-diffusion-xl-base-1.0
kvad images make "a red fox in fresh snow" \
--model stabilityai/stable-diffusion-xl-base-1.0The picture is written to image-ID.png, or to --out FILE. It is also kept
on the server with the settings that made it, where the web UI's Images page
and kvad images list it.
#Models
| Model | In memory | A step, on an M5 Pro | Steps |
|---|---|---|---|
| Stable Diffusion 1.5 | 1.9 GB | 0.75 s at 512² | 25 |
| SDXL | 6.4 GB | 3.1 s at 1024² | 30 |
| FLUX.1-schnell, q8 | 18.3 GB | about 9 s at 1024² | 4 |
| Qwen-Image, q8 | about 29 GB | 28 s at 1024² | 20 |
SDXL and Qwen-Image run the denoiser twice per step while guidance is on, and those figures include both passes. FLUX.1-schnell needs no guidance. FLUX.1-dev is not supported.
#Fine-tunes
SDXL and Stable Diffusion 1.5 fine-tunes load in any of the three layouts people publish: a diffusers folder, a single checkpoint file in Stability's own layout, or a file on your disk.
kvad pull Lykon/dreamshaper-8
kvad images make "a lighthouse at dusk" --model Lykon/dreamshaper-8
kvad images make "a lighthouse at dusk" --model ~/Downloads/some-sdxl-finetune.safetensorsA repository with several checkpoints is named as repo:file.safetensors.
#GGUF
Community GGUF files of Qwen-Image's and FLUX.1-schnell's transformers load by
the repository and the quantisation together. The rest of the model comes from
the base_model on the model card.
kvad pull city96/Qwen-Image-gguf:Q4_K_S
kvad images make "a red fox in fresh snow" --model city96/Qwen-Image-gguf:Q4_K_SA Q4_K_S holds 9.1 GB less than q8 for Qwen-Image, and 5.5 GB less for FLUX, for a step that is 5 to 18% longer.
#Options
kvad images make PROMPT [--out FILE] [--model MODEL] [--size WxH] [--steps N]
[--guidance F] [--negative TEXT] [--seed N]
[--lora NAME[:SCALE]]...Anything left out is the model's own default. --lora applies a
LoRA for this picture only.
kvad images # the gallery, as a list
kvad images rm 12#From the API
POST /v1/images/generations is OpenAI's images endpoint, with the settings
OpenAI has no field for beside its own.
curl http://127.0.0.1:5823/v1/images/generations \
-H 'content-type: application/json' \
-d '{
"model": "stabilityai/stable-diffusion-xl-base-1.0",
"prompt": "a red fox in fresh snow",
"size": "1024x1024",
"steps": 30,
"guidance_scale": 5,
"seed": 5,
"response_format": "url"
}'With "stream": true the server sends an event per denoising step, a rough
preview beside each when partial_images asks for one, and a completed event
per image. See OpenAI-compatible API.
#Memory
The VAE decode is the most memory an image asks for: it is the only stage at
full resolution. If a picture fails with a message about GPU memory, another
model is probably loaded. kvad ps shows what is, and kvad unload frees it.
#The pictures on the front page
The four pictures on the front page were made with the commands below, on an M5 Pro, with SDXL's defaults: 1024², 30 steps, guidance 5. Each took 82 to 86 seconds to denoise and 9 to decode. The same seed draws the same picture.
M=stabilityai/stable-diffusion-xl-base-1.0
kvad images make "a lighthouse on a rocky Norwegian coast at dusk, calm sea, warm light in the lamp room, long exposure photograph" --model $M --seed 11
kvad images make "a red fox in fresh snow at the edge of a birch forest, early morning light, wildlife photograph" --model $M --seed 5
kvad images make "a wooden stave church in a green valley under low clouds, watercolour" --model $M --seed 3
kvad images make "an open book on a wooden desk by a window, a single candle, oil painting, Dutch golden age" --model $M --seed 21