Serve

OpenAI-compatible API

Chat completions, images and videos in OpenAI's shapes, so existing clients and SDKs work unchanged.

The base URL is the server's address with /v1:

http://127.0.0.1:5823/v1

With the default configuration there is no authentication on loopback, and clients that insist on an API key can be given any string. With an authentication mode configured, send an API key as a bearer token.

#With an SDK

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:5823/v1", api_key="unused")

reply = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[{"role": "user", "content": "Say hello in Norwegian."}],
)
print(reply.choices[0].message.content)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "http://127.0.0.1:5823/v1", apiKey: "unused" });

const stream = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-1.5B-Instruct",
  messages: [{ role: "user", content: "Say hello in Norwegian." }],
  stream: true,
});
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");

#Chat completions

POST /v1/chat/completions

curl http://127.0.0.1:5823/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{
    "model": "Qwen/Qwen2.5-1.5B-Instruct",
    "messages": [
      {"role": "system", "content": "You are terse."},
      {"role": "user", "content": "Why is the sky blue?"}
    ],
    "stream": true
  }'

The request takes messages, stream, temperature, top_p, max_tokens (or max_completion_tokens), seed, tools and tool_choice.

stream: true returns text/event-stream in OpenAI's chunk format, ending with [DONE]. Otherwise the answer is one JSON object. The numbers for each generation, prefill and decode rates and cached tokens, are in an extension field.

#Naming a model

model is a repository, or repo@backend for a particular backend:

{ "model": "Qwen/Qwen3-14B@gpu-q4" }
The model is What happens
in memory It answers.
on disk, and fits beside what is loaded It is loaded, then answers.
on disk, and does not fit 409, naming what holds the memory. Nothing is unloaded.
not on disk 404. Pull it first.
left out The one model in memory. Refused when there are several.

#Reasoning

A reasoning model's working arrives as reasoning_content, the field DeepSeek's API established, and content is the answer alone. Both stream.

#Tools

tools are passed to the model's own chat template and its calls come back as tool_calls. See Tool calls and agents.

#Models

GET /v1/models lists every model on the machine in OpenAI's shape, with a kvad object on each:

{
  "id": "Qwen/Qwen2.5-1.5B-Instruct",
  "object": "model",
  "owned_by": "huggingface",
  "kvad": { "tools": true, "resident": false, "kind": "chat" }
}

kind is chat, image or video. tools says whether the model's template has a place for tool definitions. resident says whether it is in memory now.

#Images

POST /v1/images/generations

OpenAI's fields are prompt, model, n (at most 4), size as WIDTHxHEIGHT, and response_format of b64_json (the default) or url. Beside them are the settings OpenAI has no field for:

Field
negative_prompt
steps or num_inference_steps
guidance_scale
seed
loras [{ "name": "…", "scale": 1 }], at most 4
stream, partial_images An event per denoising step, with a preview.

Each image in the answer carries a kvad object with its id, its seed, its settings and the time each stage took.

When streaming, the events are image_generation.step, image_generation.partial_image and image_generation.completed.

#Videos

POST /v1/videos starts a job and answers at once.

Route
POST /v1/videos Start one. JSON, or multipart/form-data as OpenAI's SDKs send it.
GET /v1/videos The list, newest first.
GET /v1/videos/{id} One video, and how far along it is.
GET /v1/videos/{id}/events Follow it until it ends.
GET /v1/videos/{id}/content The file.
DELETE /v1/videos/{id} Delete it, stopping it if it is still being made.

The request takes prompt, model, size, and seconds or frames, plus fps, seed, audio, input_reference, steps, guidance_scale, negative_prompt, decoder, pipeline and loras. Video explains what each one does.

#Everything else

The server's own API is under /api: models, conversations, jobs, datasets, evals, benchmarks, metrics, accounts. It is documented by the server itself.

  • http://127.0.0.1:5823/api is the reference, as a page.
  • /api/openapi.json is the same thing as OpenAPI. It is generated from a table that a test compares against the router, so a route that exists is a route that is documented.
  • kvad api lists every route, and kvad api GET /api/health calls one.

#What is not there

  • tool_choice accepts auto and none. required and a named function are refused, because forcing a call would need a constrained sampler.
  • There is no embeddings endpoint and no legacy /v1/completions. Raw completion without a chat template is POST /api/playground/complete.
  • Requests are answered one at a time.
esc

Try , or .