Use

LoRAs

Apply published LoRAs to image and video models per request, and train one for SDXL on your own pictures.

#Applying one

LoRAs apply to Qwen-Image, FLUX, SDXL, Stable Diffusion 1.5 and LTX-2.5. They are applied at run time, per request: the model stays loaded as it was, and each request chooses its own LoRAs and strengths. Nothing is merged into the weights and nothing is reloaded.

kvad pull lightx2v/Qwen-Image-Lightning:Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors

kvad images make "a tiny astronaut hatching from an egg on the moon" \
    --model Qwen/Qwen-Image --steps 8 \
    --lora lightx2v/Qwen-Image-Lightning:Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors

A strength goes after a colon, --lora NAME:0.8, and is 1 if left out. Up to four can be applied to one request. A LoRA is named the way it was pulled: repo, repo:file.safetensors, or a path on this machine.

The formats people publish are all read: PEFT's, kohya's, diffusers' and Black Forest Labs', including text-encoder and convolution pairs.

That Lightning LoRA is worth having. With it Qwen-Image draws in 8 steps instead of 50: 68 seconds at 1024² on an M5 Pro against 448, and a sharper picture.

A LoRA that does not fit the model fails the request, and the message names the layers it adapts that the model does not have.

#Training one

A LoRA for SDXL can be trained on a folder of pictures, each with a caption in a .txt file of the same name beside it. That is the kohya and diffusers convention.

my-photos/
  01.jpg
  01.txt      "a photo of a red boat, my-style"
  02.jpg
  02.txt
kvad-gpu tune --data ./my-photos --name my-style

Training arrived after v0.11.0. Until the next release it needs a build from source.

The result is ~/.local/share/kvad/loras/my-style.safetensors, which kvad images --lora applies and so does diffusers.

kvad images make "a lighthouse, my-style" \
    --model stabilityai/stable-diffusion-xl-base-1.0 \
    --lora ~/.local/share/kvad/loras/my-style.safetensors

#What a run takes

Measured on an M5 Pro with 48 GB, one picture a step:

Size A step While training At most
512² 1.9 s 6.4 GB 7.2 GB
1024² 7.3 s 7.6 GB 13.7 GB

A thousand steps at 1024², SDXL's own size, is about two hours. At 512² it is about half an hour, and SDXL draws less well there.

#Options

Option Default
--size N 1024 Pixels a side, a multiple of 64.
--rank N 16
--alpha F the rank The LoRA is scaled by alpha / rank.
--steps N 1000
--lr F 1e-4
--eval-every N 100 Steps between validation measurements.
--holdout N a tenth Pictures kept out of training, to measure on.
--caption TEXT The caption of every picture that has none.
--from FILE Go on from a LoRA this wrote.
--cap GB ¾ of memory End the run if its memory passes this.

#Which step is kept

The training loss is not the number to watch: it keeps falling while the LoRA memorises your pictures. So some pictures are held out, the loss on those is measured every --eval-every steps, and the LoRA written to NAME.safetensors is the step where that was lowest. The last step is written beside it as NAME.last.safetensors when it is a different one.

ffmpeg is needed, to decode the pictures. --cap is there because a backward pass that does not fit in memory is not refused by macOS: it takes the machine down.

Training a LoRA has the long form, with one run described from start to finish.

esc

Try , or .