Skip to content

Inclavate REST API Reference ​

Base URL: http://<orchestrator-host>:<port> (default port: 8001)

All responses are JSON unless noted. Timestamps are RFC 3339. Layer indices are 0-based, exclusive end (start_layer..end_layer).


Authentication ​

Admin endpoints require the X-Admin-Token header:

X-Admin-Token: <token>

The token is set with --admin-token when starting the orchestrator.
Without a token configured, all admin routes return 403.


System ​

GET /health ​

Returns orchestrator liveness.

Response 200

json
{ "ok": true, "version": "0.1.0" }

Models ​

GET /v1/models ​

OpenAI-compatible model list.

Response 200

json
{
  "object": "list",
  "data": [
    { "id": "llama3.2:1b", "object": "model", "owned_by": "dppan" }
  ]
}

GET /api/models ​

Live node coverage plus multimodal capability metadata for models currently represented by registered nodes.

Response 200

json
[
  {
    "id": "gemma4:12b",
    "n_layers": 48,
    "d_model": 3840,
    "node_count": 2,
    "active_nodes": 2,
    "idle_nodes": 0,
    "dead_nodes": 0,
    "covered": [[0, 24], [24, 48]],
    "gaps": [],
    "chain_complete": true,
    "total_vram_mb": 32768,
    "input_modalities": ["text", "image", "audio", "video_frames"],
    "supports_tools": true,
    "reasoning_effort_levels": [],
    "projector_state": "ready",
    "decoder_requirements": {
      "m_rope": false,
      "non_causal_media": true
    },
    "unsupported_reason": null
  }
]

reasoning_effort_levels lists the reasoning_effort values this model accepts. It is derived when the model loads, not tabulated, and comes from one of two mechanisms:

  • Native — the model's chat template reads an effort variable, so the level is passed to it and the model was trained on the distinction. Probed by rendering, because support does not follow architecture (Qwen3.6 has none, Qwen3.8 does) and the variable is not always spelled the same (gpt-oss and Qwen3.8 read reasoning_effort; Muse-Glimmer reads reasoning_strength). Levels are whatever that model accepts.
  • Token budget — the model reasons but its template exposes no control at all (Kimi-K2-Thinking, DeepSeek-R1). The level becomes a cap on how many tokens the reasoning span may use, after which its closing tag is forced. Levels are always ["low", "medium", "high", "max"] — 512, 2048, 8192 tokens and no limit. Requires the template to declare reasoning tags to close.

Native wins when a model has both, so the list never mixes the two. A model with neither returns [] and has no effort control; render no selector for those rather than a disabled one.

"none" is accepted by the chat endpoints on every model and is deliberately absent from this list — it is applied as "reasoning off" and never reaches a template, so probing it would report it unsupported on the models where it is the one value that always works.

GET /models/registry ​

Authoritative catalog plus local config/models.toml overlay. Unlike GET /api/models, this includes models with no currently registered nodes.

Response 200

json
[
  {
    "model_id": "gemma4:12b",
    "n_layers": 48,
    "d_model": 3840,
    "gguf_sha256": "",
    "quant_type": "Q4_0",
    "arch": "gemma4",
    "engine_loaded": false,
    "source": "catalog",
    "input_modalities": ["text", "image", "audio", "video_frames"],
    "projector_state": "ready",
    "decoder_requirements": {
      "m_rope": null,
      "non_causal_media": null
    },
    "unsupported_reason": null
  },
  {
    "model_id": "llama3.2:1b",
    "n_layers": 16,
    "d_model": 2048,
    "gguf_sha256": "74701a8c...",
    "quant_type": "Q4_K_M",
    "arch": "llama",
    "engine_loaded": true,
    "source": "local",
    "input_modalities": ["text"],
    "projector_state": "missing",
    "decoder_requirements": {
      "m_rope": null,
      "non_causal_media": null
    },
    "unsupported_reason": null
  }
]

projector_state is ready, missing, or invalid. A valid vision projector also advertises video_frames. Decoder requirement values are null before the first lazy projector load and become booleans after libmtmd derives them; they are never guessed from model names. unsupported_reason is populated when a discovered sidecar fails metadata validation, inspection, or decoder compatibility loading.

GET /admin/chain/:model_id (admin) ​

Chain health for a specific model — shows which layer ranges are covered.

Response 200

json
{
  "model_id": "llama3.2:1b",
  "n_layers": 16,
  "covered": [[0, 8], [8, 16]],
  "gaps": [],
  "ready": true
}

Chat Completions ​

POST /v1/chat/completions ​

OpenAI-compatible chat endpoint (non-streaming and SSE streaming).

Request body

json
{
  "model": "llama3.2:1b",
  "messages": [
    { "role": "system",    "content": "You are a helpful assistant." },
    { "role": "user",      "content": "What is 2 + 2?" }
  ],
  "max_tokens":   200,
  "temperature":  0.8,
  "top_p":        0.95,
  "stream":       false
}
FieldTypeDefaultDescription
modelstringrequiredModel ID from the registry
messagesarrayrequiredConversation; role is system, user, or assistant
max_tokensinteger512Max new tokens to generate
temperaturefloat0.8Sampling temperature (0 = greedy)
top_pfloat0.95Nucleus sampling p
thinkingbooleanmodel defaultForce reasoning on or off
reasoning_effortstringmodel defaultHow hard to think — see below
streambooleanfalseEnable SSE streaming

reasoning_effort ​

OpenAI's spelling. "none" turns reasoning off and works on every model. Any other value must be one of the model's reasoning_effort_levels from GET /api/models; anything else is a 400 listing what the model accepts.

How a level is applied depends on the model, and the two mechanisms are described under GET /api/models above. On a native model the level reaches the model's own chat template, so it means whatever that model was trained to make it mean. On a token-budget model it caps the reasoning span instead: low 512 tokens, medium 2048, high 8192, max no limit. When the cap is reached the model's own reasoning end tag is forced, so the reply still parses normally — it just stops deliberating and answers.

The rejection is deliberate. A template that does not recognise a value may raise rather than ignore it — Qwen3.8 raises on OpenAI's own minimal and max — and a raised template silently falls back to a plainer prompt with none of the model's reasoning instructions. Failing the request is the only way that surfaces.

A request that omits the field renders exactly as it did before the field existed. thinking and reasoning_effort are independent: the first decides whether the model reasons, the second how hard, and "none" sets the first.

The same field is accepted on the /ws/chat request frame. POST /v1/responses does not accept it yet.

messages[].content may also be an ordered array of multimodal parts. Plain string content remains fully backward-compatible.

json
{
  "model": "gemma4:12b",
  "messages": [{
    "role": "user",
    "content": [
      { "type": "text", "text": "Compare these frames." },
      {
        "type": "image_url",
        "image_url": { "url": "data:image/png;base64,<base64>" }
      },
      {
        "type": "input_audio",
        "input_audio": { "data": "<base64>", "format": "wav" }
      },
      {
        "type": "input_video",
        "input_video": {
          "frames": [
            {
              "image_url": { "url": "data:image/jpeg;base64,<base64>" },
              "timestamp_ms": 0
            },
            {
              "image_url": { "url": "data:image/jpeg;base64,<base64>" },
              "timestamp_ms": 1000
            }
          ]
        }
      }
    ]
  }]
}

Video: frames or a whole file ​

input_video takes either frames or file, never both. The example above sends frames the client extracted itself. To hand the server a whole video and let it decode the frames:

json
{
  "type": "input_video",
  "input_video": { "file": "data:video/mp4;base64,<base64>" }
}

file is the data URL itself, not an object. Accepted containers are video/mp4, video/quicktime, video/webm, and video/x-matroska; the declared MIME type must match the container's magic bytes.

Server-side decoding requires ffmpeg and ffprobe on the host's PATH. Deployments without them report supports_video_files: false on /api/models and reject file with an explanatory error — check that field before uploading. Frame lists work regardless. The server decodes at most 240 frames from one file.

Generation from a chat turn ​

A model can produce audio mid-conversation by calling a generation tool. Those tools are opt-in: they are offered only when the request asks for them with tool_choice ("auto", "required", or a named function). Without it an ordinary chat turn is unchanged, and cannot hand a probabilistic model the authority to start expensive work nobody requested.

When a reply starts one, the response carries a generated array of job references alongside the usual content, and the audio is fetched the same way as any other generation.

Three optional fields steer what the model produces. Each is the caller's choice rather than the model's — it has no way to know a user's budget or which voice they picked in a UI:

Field
generation_modelRestrict generation to one model. Naming an unconfigured one offers no tools rather than falling back.
generation_voiceRestrict it to one voice, likewise.
generation_max_duration_secondsCeiling on the audio a reply may produce.

A reply may start at most four generations; anything beyond that is handed back as an ordinary tool call rather than silently dropped.

Limits ​

Supported image MIME types are PNG, JPEG, WebP, and BMP. WebP is decoded and re-encoded as PNG before it reaches the model, so its byte budget is counted against the PNG, which is usually larger than the WebP you sent. Audio format is wav or mp3. Video frames must be in nondecreasing timestamp order. Remote HTTP URLs and local paths are disabled; images, video frames, and video files must use base64 data URLs. Requests are limited to 24 MiB, 64 content parts, 16 media items, 16 frames per video part, 8 MiB encoded per item, 16 MiB encoded total, 16 megapixels per image, 32 megapixels total, and five minutes per audio item. HTTP streaming, HTTP non-streaming, and /ws/chat use the same preparation and validation path.

Response 200 (non-streaming)

json
{
  "id": "chatcmpl-<uuid>",
  "object": "chat.completion",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "4" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 1,
    "total_tokens": 13
  }
}

When the model reasoned, its thinking comes back beside the answer rather than inside it:

json
"message": {
  "role": "assistant",
  "reasoning_content": "The user wants 2+2. That is 4.",
  "content": "4"
}

reasoning_content is absent when the model did not reason, so a client that ignores unknown fields always finds the answer alone in content. The name is the one vLLM, SGLang and DeepSeek use; OpenAI's chat schema has no equivalent field, and reasoning tokens are counted in completion_tokens like any other.

Reasoning is spent from the same max_tokens budget as the answer. A model that thinks for longer than the budget allows returns its reasoning with an empty content — raise MAX_TOKENS (default 512) or lower the effort level. finish_reason is "length" when that happens, which is how a client tells it apart from a model that simply had little to say.

finish_reason takes three values, on both the non-streaming response and the final SSE frame:

ValueMeaning
stopThe model finished on its own or hit a stop sequence
lengthmax_tokens was reached; the reply is cut off
tool_callsThe reply is a tool call for the caller to run

length outranks tool_calls: a reply stopped at the cap may have been cut mid-call, so the truncation is reported instead.

Response 200 (streaming, Content-Type: text/event-stream)

Each event:

data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}

data: [DONE]

Errors

CodeMeaning
503No Ready nodes covering all layers of the requested model (chain gap)
422Model not in registry
400Malformed request

Media validation and capability failures return a stable JSON envelope:

json
{
  "error": {
    "message": "unsupported modality: ...",
    "type": "unsupported_modality",
    "param": null,
    "code": "unsupported_modality"
  }
}

DELETE /sessions/:id ​

Close a generation session and free KV cache on all nodes.

Response 204 — no body


Artifacts ​

Durable blobs the orchestrator keeps beyond a single request — an upload you reference later by ID, or bytes a generator produced. Bytes never reach decoder nodes. See Media for the storage model and configuration.

Ownership follows the deployment: with no database every request is the single local principal; with a database the owner is the session's account, and an unauthenticated request is 401 rather than a fallback. A missing artifact and another owner's artifact both return 404.

If artifact storage is not configured, every endpoint here returns 503.

POST /v1/artifacts ​

Upload an artifact. The body is streamed; Content-Type is recorded as the media type and defaults to application/octet-stream. A single upload is capped at 512 MiB.

Response 201

json
{
  "id": "1deaf021-07a3-407c-9d6f-1af47c23db18",
  "media_type": "audio/wav",
  "size_bytes": 40000,
  "digest": "5327345f0dc6d1c5dc69b1edbd486188bf47ee65bf3ecb4e6168e3072cc45b1b",
  "state": "ready",
  "created_at": "2026-08-06T09:14:22.981Z",
  "expires_at": null
}

digest is SHA-256 of the full object. state is pending, ready, or failed; only ready artifacts are readable.

Errors — 413 body over the per-request cap · 507 owner quota exhausted

GET /v1/artifacts ​

List the caller's own artifacts, newest first.

Query — limit (default 100, max 1000) · offset (default 0) · kind (generated | uploaded; absent lists both)

kind splits the store by origin, which each artifact records at creation: uploaded for anything a person sent — through either the Files API or POST /v1/artifacts — and generated for what a job produced. Both are audio/wav, and nothing else about them differs.

Origin is stored rather than inferred from filename. Only the Files API records a filename, so reading the split off it answered "which endpoint made this" instead of "what is this", and filed a clip uploaded through POST /v1/artifacts as generated output.

The filter is applied in the query rather than left to the client because the two halves are managed separately — generated audio is reproducible, an uploaded clip may be the only copy in existence. A client that fetched a mixed page and dropped the rows it did not want would get pages of unpredictable size, a "next" button it could not compute, and totals that disagreed with the rows. An unrecognised value lists both rather than erroring: showing more than asked is recoverable by the reader; an empty page for a typo looks like an empty store.

Response 200

json
{
  "artifacts": [
    { "id": "1deaf021-...", "media_type": "audio/wav", "size_bytes": 40000,
      "digest": "5327345f...", "state": "ready",
      "created_at": "2026-08-06T09:14:22.981Z", "expires_at": null,
      "filename": "my-voice.wav", "purpose": "user_data" }
  ],
  "truncated": false
}

filename and purpose are present only for artifacts uploaded through the Files API; anything created natively has neither. They are display metadata, not the origin discriminator — use kind for that.

truncated says the server had more than it returned — walk the rest with offset. Use GET /v1/artifacts/usage for totals: summing pages understates them, and is wrong the moment anything is deleted mid-walk.

Ordering is (created_at DESC, id DESC). The tie-break on id is what makes offset safe: two artifacts written in the same millisecond would otherwise be free to swap places between requests, showing one twice and hiding the other.

GET /v1/artifacts/usage ​

How much this caller is storing, and what it is measured against.

Response 200

json
{ "bytes": 43219872, "objects": 37, "max_bytes": 0, "max_objects": 0,
  "generated": { "bytes": 42100000, "objects": 34 },
  "uploaded":  { "bytes":  1119872, "objects":  3 } }

generated and uploaded are the same totals split by origin, and they sum to the top-level pair. They are here so each half of a management view can state its own size without walking its pages — and so a bulk delete can say what it will free before it starts, which is the difference between "delete everything" and "delete all 34 files (42.1 MB)".

max_bytes and max_objects are 0 when no ceiling is configured, which is the self-hosted default. default_ttl_hours is present only when DPPAN_ARTIFACT_TTL_HOURS is set; its absence means artifacts are kept until deleted explicitly and nothing is swept.

Counted with a COUNT/SUM over the whole account rather than by summing a listing, so it stays correct past the list limit.

GET /v1/artifacts/:id ​

Stream the bytes. Responds with Accept-Ranges: bytes.

Supply Range: bytes=<start>-<end> for a partial read; the response is 206 with a Content-Range header. Single ranges only — multi-range (bytes=0-9,20-29) and suffix (bytes=-500) requests return 416 rather than being partially honoured.

Errors — 400 malformed ID or range · 404 unknown or not yours · 409 artifact is not ready · 416 range outside the object

GET /v1/artifacts/:id/meta ​

Metadata only; same body shape as the upload response.

DELETE /v1/artifacts/:id ​

Remove bytes and metadata.

Response 204 — no body. 404 if it does not exist or is not yours.


Files (OpenAI-compatible) ​

An adapter over Artifacts, not a second store. A file is an artifact: same IDs, same ownership, same quota, same retention — so a clip uploaded here can be referenced from DPPAN-native endpoints, and the reverse. Point an OpenAI client at this base URL and files.create() works.

MethodPath
POST/v1/filesMultipart upload. Fields: purpose, file. Returns 201.
GET/v1/filesList your own files.
GET/v1/files/:idMetadata.
GET/v1/files/:id/contentThe bytes.
DELETE/v1/files/:idDelete.
bash
curl -X POST http://localhost:8001/v1/files \
  -F purpose=user_data -F file=@my-voice.wav
# {"id":"66815206-...","object":"file","bytes":240044,
#  "filename":"my-voice.wav","purpose":"user_data"}

The returned id is an artifact ID, so it can be passed straight to a generation as speaker_artifact.

Where compatibility stops

purpose is validated rather than accepted. assistants, user_data and vision work; fine-tune, batch and evals describe workflows this deployment does not run and are refused by name rather than stored — accepting one would report success for something that will never happen.

Uploads are capped at 64 MiB. Multipart arrives whole rather than streaming, so the memory cost of an upload is roughly twice its size; a limit this path could not honour safely would not be a limit.

Downloads are always served as an attachment with nosniff, and a media type outside a known-inert set is served as application/octet-stream. The type is chosen by whoever uploaded the file, and a document served under its own type from this origin would run against the viewer's session.


Generation ​

Asynchronous media generation. Submit a job, poll it, then download the artifacts it produced. See Media for the model.

Ownership follows the same rule as artifacts: with no database every request is the single local principal; with a database the owner is the session's account. A missing job and another owner's job both return 404.

If generation is not configured, every endpoint here returns 503.

Which modalities work depends on what is configured. Speech is available when a speech model has been pulled with dppan pull --speech, or named explicitly with DPPAN_SPEECH_MODEL and DPPAN_SPEECH_MMPROJ (see Media); image and video have no generator, so they answer 422.

Rather than inferring what a deployment supports, ask it — GET /v1/generations/capabilities.

GET /v1/generations/capabilities ​

What this deployment can generate, and what each generator will honour. Answer this before composing a request: a client that assumes limits offers controls the deployment silently clamps.

Unauthenticated on purpose — it exposes only which generators are configured and reveals nothing about anyone's jobs.

json
{
  "generators": [
    {
      "modality": "audio",
      "model": "qwen3-tts",
      "voices": ["warm", "clear"],
      "requires_speaker": true,
      "languages": ["de", "en", "es", "fr", "it", "ja", "ko", "pt", "ru", "zh"],
      "limits": {
        "max_duration_seconds": 40.96,
        "default_temperature": 0.9,
        "default_top_k": 50,
        "default_top_p": 1.0,
        "default_repetition_penalty": 1.05
      }
    },
    {
      "modality": "audio",
      "model": "hf-PkmX-orpheus-3b-0.1-ft-Q8_0-GGUF-Q8_0-full",
      "builtin_voices": ["tara", "leah", "jess", "leo", "dan", "mia", "zac", "zoe"],
      "requires_speaker": false,
      "limits": {
        "max_duration_seconds": 43.69,
        "default_temperature": 0.9,
        "default_top_k": 50,
        "default_top_p": 1.0,
        "default_repetition_penalty": 1.05
      }
    },
    {
      "modality": "audio",
      "model": "ai4bharat-IndicF5",
      "requires_speaker": true,
      "requires_reference_text": true,
      "speed_range": { "min": 0.3, "max": 2.0 },
      "languages": ["as", "bn", "gu", "hi", "kn", "ml", "mr", "or", "pa", "ta", "te"]
    }
  ]
}

voices, builtin_voices and languages are omitted when empty, requires_reference_text when false, and speed_range when the generator has no pace control.

voices and builtin_voices are both passed as voice, but they are not the same thing. voices are reference clips this deployment registered, and an operator can add or remove them; builtin_voices are speakers baked into the checkpoint's weights, and nobody can. Offer them under one control — the server knows which is which — but do not merge them in your own configuration, because only one of the two is yours to change.

requires_speaker says whether a request must carry a voice or speaker_artifact at all. It is a property of the checkpoint, not of speech or even of the codec — two Qwen3-TTS models on the same codec answer differently. Qwen3-TTS Base carries a speaker encoder and produces unreliable length and level when asked to speak unconditioned, so a clip is required; Qwen3-TTS CustomVoice has no encoder at all and offers nine builtin_voices instead; Orpheus has no clip conditioning whatsoever, and sending it one is refused with speaker_unsupported. Gate your submit control on this field: a client that assumes either answer disables the control for half the models it can see.

requires_reference_text says the clip must come with exactly what is said in it, as reference_text. It is true for F5 models (IndicF5), which continue the clip's speech rather than encoding a voice from it: a clip without its transcript is refused with 400, and so is a configured voice, which has no transcript. Every other speech model refuses reference_text with 400, so send it only where this field asks for it. An F5 generator reports no limits — it has no sampler and no frame ceiling to state; max_duration_seconds on the request still caps the result. See Media → speech for what makes a good reference clip.

speed_range says the generator honours speed — in params, or OpenAI's speed on /v1/audio/speech — and within what bounds. It divides the generated duration: 2.0 is twice as fast, 0.5 half. A value outside the range is refused with 400. Show a speed control only where this is present: every other generator ignores the field, as it always has.

languages is read from the model's own vocabulary, so it is what this model can actually speak rather than what the library could translate; anything outside it fails the job. Empty means the model has no language tokens to select between — an Orpheus checkpoint is fine-tuned for one language, fixed when it was pulled — not that it speaks nothing, so render it as "no choice to make" rather than as an empty dropdown.

limits.max_duration_seconds is the longest single generation the frame budget allows, and differs between generators on the same budget because the frame rate is the codec's own. It is not a limit on how much you can have spoken: longer text is split and joined, so the finished audio can be many times this. To cap the result, send max_duration_seconds on the request — that one bounds the finished audio.

An empty generators list is the honest answer for a deployment with no generator, and a client should render it as such rather than offering a form whose submission cannot succeed.

POST /v1/generations ​

Body

json
{
  "model": "qwen3-tts",
  "modality": "audio",
  "params": { "text": "Hello world" },
  "seed": 42,
  "input_artifacts": ["1deaf021-..."],
  "idempotency_key": "submit-1",
  "deadline_seconds": 120
}

modality is audio, image, video, or text. text runs the arrow the other way — media in, an answer out — and is documented under Understanding below. params is generator-specific and passed through untouched. input_artifacts are artifact IDs consumed as input — a speaker reference for voice cloning, an image to animate, the clip an understanding job reads. Submitting the same idempotency_key twice returns the first job rather than queueing a second. deadline_seconds may lower the server default but not exceed 24 hours.

Response 202 — the job is durably queued, not finished.

json
{
  "id": "9f1c...", "state": "queued", "modality": "audio", "model": "qwen3-tts",
  "seed": 42, "input_artifacts": [], "output_artifacts": [],
  "progress_stage": null, "progress_fraction": null,
  "error_code": null, "error_message": null,
  "created_at": "2026-08-06T09:14:22.981Z", "started_at": null,
  "completed_at": null, "expires_at": "2026-08-06T09:24:22.981Z"
}

Errors — 400 malformed request · 409 idempotency-key race · 422 no generator produces this · 429 queue full

Those last two are worth telling apart: 429 is retryable, 422 never will be.

Understanding: media in, text out ​

Every other generator turns a prompt into bytes. modality: "text" turns bytes into a prompt's answer — describe this photo, transcribe this clip, read this receipt. It is the same queue, the same job record and the same artifact store; only the direction changes.

It is worth knowing when to use this instead of a multimodal chat completion, which reads media too. Chat streams its answer and needs a connection held open. A job does not: it survives a restart, it can be polled or subscribed to later, and it is the right shape for a batch of receipts or a clip that takes a minute to read.

Discovering them. Understanding generators appear in GET /v1/generations/capabilities like any other, with kind: "understand":

json
{
  "modality": "text",
  "kind": "understand",
  "model": "qwen3vl:4b",
  "ready": true,
  "input_modalities": ["image", "video"]
}

input_modalities is what this model's projector handles, so a file picker can refuse a clip before it becomes a job that fails on load. An OCR checkpoint lists image and no audio; Ultravox lists audio and no image.

video appears only when the orchestrator host has ffmpeg on its PATH — the whole container is decoded server-side, because a job has no browser to extract frames the way the chat path does. A model that reads video on a host without ffmpeg simply does not advertise it.

ready says whether the model can serve right now. A model can be installed, multimodal and projector-ready and still be unservable, because nothing has loaded it or the nodes covering its layers are not all connected. Render an unready generator as visible-but-unselectable rather than hiding it: "that model exists and is not up" is a different problem from "not installed", and only one of them is fixed by pulling a model. voices, languages and limits are absent — understanding has no voice to pick and no duration budget.

Postgres deployments need migration 0009 first

text was added to the output_modality CHECK constraint by rust/migrations/0009_generation_job_text_modality.sql, and nothing in the binary applies it — hosted migrations are run by an operator on purpose. Until it is applied, every understanding job fails at insert with a 500 and a raw constraint-violation message. Self-hosted SQLite deployments are unaffected: that schema ladder applies itself on start.

Submitting one. The media goes in input_artifacts; the question goes in params.prompt. Upload first (POST /v1/files or POST /v1/artifacts), then:

bash
curl -X POST http://localhost:8001/v1/generations \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3vl:4b","modality":"text",
       "input_artifacts":["66815206-..."],
       "params":{"prompt":"What is the total on this receipt?","max_tokens":512}}'
params field
promptrequiredWhat to look for. An empty or absent prompt fails the job rather than guessing — no_prompt
max_tokens1024Longest answer to generate
thinkingmodel defaulttrue or false. Absent leaves the model's own setting. Worth sending false to a reasoning model asked to read a receipt, which will otherwise spend the whole budget deciding how to read it

More than one input artifact is allowed and order is preserved — the media is sent in the order given, with the prompt last, which is the order these templates render. Each input is capped at 64 MiB.

Response. A 202 with the usual job record, modality reading "text". When it completes, output_artifacts[0] is a single text/plain artifact (capped at 4 MiB) that you fetch from GET /v1/artifacts/:id. Progress reports three stages — reading input, understanding, publishing — over GET /v1/generations/:id/events.

The artifact is the raw token stream

Nothing parses it on the way out, so a reasoning model's chain-of-thought arrives inside it — <think>…</think>, or gpt-oss harmony's <|channel|>analysis<|message|>…<|channel|>final<|message|>. Split on those markers before displaying the answer, or send "thinking": false. A client that prints the artifact verbatim shows the user the model's working.

Job errors — reported in error_code on the job record, not as an HTTP status, because they happen after the 202:

code
no_promptparams.prompt was empty or absent
no_inputno input_artifacts were given
input_unreadablean input artifact is gone, or its bytes could not be read
input_too_largean input exceeds 64 MiB
unsupported_inputthis model cannot read that media type, or a video was sent to a host with no ffmpeg
inference_failedthe model is not loaded, its node chain is incomplete, or generation failed
publish_failedthe answer could not be written to the artifact store

GET /v1/generations ​

The caller's own jobs, newest first. ?limit= (default 100, max 1000), ?offset= (default 0).

output_artifacts here lists only artifacts that still exist. A job record outlives the bytes it produced — artifacts expire, and they can be deleted directly — so this endpoint resolves each one against the store and reports what is actually left, with artifacts_deleted counting what is not:

json
{ "id": "…", "state": "completed", "output_artifacts": [], "artifacts_deleted": 1,
  "params": { "text": "hello world" } }

That row is not a failure: error_code is null, the prompt is intact, and it can be regenerated. Distinguishing it from a job that never produced anything is the whole reason the field exists — a client that shows a player for a completed job without checking will render a dead control. The dashboard hides these rows on the Generate page and surfaces them under Admin → Storage, where they can be cleared.

Resolution happens on the LISTING only. GET /v1/generations/:id is polled once a second by every open tab, and a store lookup there would multiply the cost of watching a job by the number of watchers to answer a question that only matters once it has finished — so a single job's artifacts_deleted is always 0.

GET /v1/generations/:id ​

Poll one job. state is queued, running, completed, failed, cancelled, or expired; only completed populates output_artifacts, which are ordinary artifact IDs fetched through GET /v1/artifacts/:id.

GET /v1/generations/:id/events ​

Live progress as Server-Sent Events, instead of polling. Each data: line is the same JSON GET /v1/generations/:id returns, sent when something changes; the stream ends with data: [DONE].

bash
curl -N http://localhost:9000/v1/generations/$JOB/events
data: {"state":"queued", ...}
data: {"state":"running","progress_stage":"loading","progress_fraction":0.0, ...}
data: {"state":"running","progress_stage":"generating","progress_fraction":0.57, ...}
data: {"state":"completed","progress_stage":"publishing","progress_fraction":1.0, ...}
data: [DONE]

The current state is sent immediately on connect, so attaching to a job that already finished returns its outcome and closes rather than hanging.

Disconnecting does not cancel the job. Watching and controlling are separate: closing a tab or losing a connection leaves the work running, and you can re-attach later or fall back to GET /v1/generations/:id. Use DELETE to actually stop it.

Any peer can serve any job's stream — progress is read from the shared job store, not from an in-process subscription, so it does not matter which orchestrator is running the job.

Errors — 400 malformed id · 404 unknown or not yours · 503 generation not configured. These arrive as status codes before the stream opens, never as an event payload.

DELETE /v1/generations/:id ​

Cancel. Returns the job rather than 204, so the caller sees which state it settled in without polling again — a cancel racing a finishing job may land as completed.

Cancellation reaches a running generator: generators stop between frames rather than being killed, so nothing is left half-written.

Errors — 404 unknown or not yours · 409 already finished (the job is still readable via GET, which is why this is not a 404)

DELETE /admin/generations/:id (admin) ​

Destroy a job record, and by default the artifacts it produced.

Admin-gated while DELETE /v1/generations/:id on the same resource is not, and the asymmetry is the point: cancelling stops work the caller started and leaves the record; this erases the record itself — the prompt, the parameters, the fact that it ever ran. Nothing else in the API can do that, and nothing undoes it.

Query — artifacts (default true). artifacts=false keeps the bytes and removes only the record, which is the exact inverse of DELETE /v1/artifacts/:id. Between the two, either half of a generation can be dropped alone.

bash
curl -X DELETE -H "X-Admin-Token: $TOKEN" \
  "http://localhost:9000/admin/generations/$JOB?artifacts=false"

Response 200

json
{ "id": "441b638c-…", "artifacts_deleted": 1 }

A missing artifact is not an error — the caller asked for it gone and it is gone — so artifacts_deleted can be lower than the job's original output count.

Errors — 401 bad or missing token · 404 unknown or not yours · 409 the job is still queued or running (a worker is writing to that record; cancel it first) · 503 no admin token configured on this deployment · 500artifacts=true on a deployment with no artifact store, which is refused rather than silently reporting the bytes deleted


Audio (OpenAI-compatible) ​

POST /v1/audio/speech ​

A thin adapter over /v1/generations — same queue, same ownership, same limits. Point an OpenAI client at this base URL and audio.speech.create() works.

Body

json
{ "model": "qwen3-tts", "input": "Hello world", "voice": "warm" }

model and input are required. Whether a speaker source is also required depends on the model — requires_speaker in capabilities answers it. Qwen3-TTS Base needs one: either a configured voice or speaker_artifact, the ID of an uploaded clip to clone. Orpheus and Qwen3-TTS CustomVoice need none, and refuse a clip if sent one; their voice names a speaker built into the checkpoint. IndicF5 needs speaker_artifact and reference_text, the clip's exact transcript — requires_reference_text in capabilities says so. language (a hint the model may ignore), instruct (a sentence of style direction, which only the instruction-tuned checkpoints act on) and max_duration_seconds (a ceiling on the finished audio, however long the text is) are the other DPPAN extensions.

Response 200 — the audio itself, not JSON.

content-type: audio/wav
x-dppan-artifact-id: 34c7dd29-...
x-dppan-job-id: 441b638c-...

The artifact outlives the response, so you can re-fetch it from /v1/artifacts/:id instead of regenerating.

Where compatibility stops

FieldBehaviour
response_formatOnly wav. Asking for mp3/opus/aac/flac returns 400 rather than WAV bytes labelled as mp3 — a client trusting the Content-Type would otherwise write a file that does not play. Omitting it gives wav.
voiceWorks, but names a voice this deployment configured or one built into the checkpoint — not OpenAI's alloy/echo/…, which do not exist here. An unknown name is refused, listing what does exist. With neither kind available there is nothing to select and any value is refused. GET /v1/generations/capabilities lists both.
speedHonoured by a generator that declares a speed_range in capabilities — IndicF5, 0.3–2.0 — where it divides the generated duration (2.0 is twice as fast); outside the range is a 400. Every other generator has no duration control and no resampling in this build, so for them it is accepted and ignored. max_duration_seconds truncates the finished audio, which is a different thing from speaking it faster.

Errors — 400 bad request or unavailable format · 422 no generator · 429 queue full · 503 generation or artifact storage not configured · 504 the wait ran out

The 504 is not a lost request: the job keeps running and the message carries its ID, so poll GET /v1/generations/:id. The wait is bounded by DPPAN_SPEECH_SYNC_TIMEOUT_SECS (default 120). Anything genuinely long-running belongs on the native async API — this endpoint is a convenience over it.

POST /v1/audio/music ​

Generate a song. Same queue and same ownership as speech, and like speech it is an adapter over /v1/generations.

Body

json
{
  "model": "minimax-music3",
  "input": "warm lo-fi hip hop, dusty piano, vinyl crackle",
  "lyrics": "[verse]\nlate night on an empty street\n",
  "duration_seconds": 30
}
Field
modelrequired — the generator's name, from capabilities. This is the bundle directory's name, e.g. minimax-music3, not the pipeline
inputrequired — the style prompt: genre, instruments, mood, tempo
lyricsoptional for most models; required by some, which need [verse] / [chorus] headers to structure the song. Omit it for an instrumental
duration_secondsclamped to the deployment's ceiling rather than refused
stepssolver steps. Higher is slower and cleaner; the default is the one the model was tuned for
stemsmix (default), vocal, or instrumental. Only some models can separate — supports_stems in capabilities says which, and the rest refuse anything but mix
response_formatwav only

Response 200 — the audio itself, with the same x-dppan-artifact-id / x-dppan-job-id headers as speech, and the same re-fetch from /v1/artifacts/:id.

Sample rate and channel count are properties of the model, not of this API — read them from the WAV header rather than assuming; supported models differ.

Expect a 504, and treat it as normal. A song is minutes of compute, so the bounded wait will usually run out and hand back a job id. That is not a failed request — the job keeps running, and GET /v1/generations/:id reports its progress. DPPAN_MUSIC_SYNC_TIMEOUT_SECS bounds the wait. For anything but a short clip, use POST /v1/generations directly rather than holding a connection open.

Progress arrives as a named phase, not a percentage — a song writes, then solves, then decodes, and those cost very different amounts, so a single fraction would sit near the end for most of the wall clock. Show the phase name.

Errors — 400 bad request · 422 no music generator configured, or the model cannot do what was asked (stems it does not separate; lyrics it requires and did not get) · 429 queue full · 503 generation or artifact storage not configured · 504 the wait ran out


POST /v1/audio/transcriptions ​

Turn a clip into text. Multipart, as OpenAI specifies, and an adapter over the same understanding job the dashboard uses — so the decode runs on the node chain like any other inference, not on the orchestrator.

Body — multipart/form-data

bash
curl -X POST http://localhost:8000/v1/audio/transcriptions \
  -F "file=@meeting.wav;type=audio/wav" \
  -F "model=qwen3-asr:1.7b"
Field
filerequired — the clip. Send a Content-Type on the part; it is what the container is decided from, and a part without one is refused
modelrequired — an understanding generator whose input_modalities include audio, from capabilities
response_formatjson (default) or text
promptoptional. A dedicated ASR model ignores it; it is passed through for models that read it
languagea hallucination filter, not a hint — a transcript the model reports as a different language comes back empty. Omit to keep whatever it detects
temperatureaccepted and ignored — transcription runs greedy

Audio formats — audio/wav, audio/x-wav, audio/wave, audio/vnd.wave, audio/mpeg, audio/mp3. Anything else is refused before the upload rather than accepted and failed inside a job. Note that browser MediaRecorder output is audio/webm or audio/mp4, neither of which is on that list — encode to WAV client-side.

Response 200

json
{ "text": "With this, Matthew's companion stopped talking." }

response_format=text returns the bare transcript as text/plain. Both carry x-dppan-artifact-id and x-dppan-job-id, so a caller can re-fetch the transcript from /v1/artifacts/:id without transcribing again.

verbose_json is refused, deliberately. It promises language, duration and per-segment timestamps. This path produces a flat transcript and has none of them, and returning the envelope with those fields missing or invented is worse than an error — a client keying on segments cannot tell the difference. srt and vtt are refused for the same reason.

Errors — 400 bad request (no model, no file, a container that cannot be decoded, an unsupported response_format) · 413 clip over 64 MiB · 429 queue full · 503 generation or artifact storage not configured · 504 the wait ran out, with the job id to poll


Nodes ​

GET /nodes ​

List all registered nodes.

Response 200 — array of NodeRecord:

json
[
  {
    "node_id":    "550e8400-e29b-41d4-a716-446655440000",
    "model_id":   "llama3.2:1b",
    "http_url":   "http://192.168.1.20:5001",
    "grpc_url":   "http://192.168.1.20:5002",
    "start_layer": 0,
    "end_layer":   8,
    "state":      "Ready",
    "vram_mb":    8192,
    "tokens_total": 1024,
    "requests_served": 12,
    "avg_latency_ms": 45.2,
    "uptime_s":   3600
  }
]

state values: Ready · Degraded · Dead · Joining

GET /nodes/rankings ​

Nodes ranked by combined score (throughput, latency, VRAM).

Response 200

json
{
  "rankings": [
    { "node_id": "...", "model_id": "llama3.2:1b", "score": 0.92, "rank": 1 }
  ]
}

The formula weights throughput against latency, and both weights are tunable on the orchestrator:

VariableDefault
DPPAN_RANK_TPS_WEIGHT1.0Multiplier on throughput (the numerator). Raise it to favour faster nodes more aggressively
DPPAN_RANK_LATENCY_WEIGHT1.0Multiplier on latency (the denominator). Set below 1.0 to de-emphasise latency

Both default to 1.0, so out of the box throughput and latency carry equal weight. They affect ranking only — a node that ranks last is still eligible for work.

POST /nodes ​

Explicit node registration with caller-provided layer range. Used by legacy (--legacy) mode.

Request body

json
{
  "node_id":    "550e8400-e29b-41d4-a716-446655440000",
  "model_id":   "llama3.2:1b",
  "http_url":   "http://192.168.1.20:5001",
  "grpc_url":   "http://192.168.1.20:5002",
  "start_layer": 0,
  "end_layer":   8,
  "n_total_layers": 16,
  "d_model":    2048,
  "vram_mb":    8192,
  "hardware":   { "gpu": "metal", "device_name": "Apple M4 Pro", "vram_mb": 32768 },
  "capabilities": [
    { "model_id": "llama3.2:1b", "n_layers": 16, "d_model": 2048 }
  ]
}

Response 200

json
{ "node_id": "550e8400-e29b-41d4-a716-446655440000", "warnings": [] }

POST /nodes/join ​

Auto-registration: orchestrator assigns the layer range based on available VRAM.

Request body

json
{
  "node_id":   "550e8400-e29b-41d4-a716-446655440000",
  "model_id":  "llama3.2:1b",
  "http_url":  "http://192.168.1.20:5001",
  "grpc_url":  "http://192.168.1.20:5002",
  "n_total_layers": 16,
  "d_model":   2048,
  "vram_mb":   8192,
  "layers":    "0-8",
  "hardware":  { "gpu": "metal", "device_name": "Apple M4 Pro", "vram_mb": 32768 },
  "capabilities": []
}

layers is optional — omit to let the orchestrator assign.

Response 200

json
{
  "node_id":    "550e8400-e29b-41d4-a716-446655440000",
  "start_layer": 0,
  "end_layer":   8,
  "warnings":   []
}
CodeMeaning
409node_id already registered
422Architecture mismatch or unknown model
503No layer range available (all layers covered)

POST /nodes/:id/health ​

Node heartbeat — updates metrics counters.

Request body (all fields optional)

json
{
  "tokens_generated": 100,
  "requests_served":  5,
  "errors":           0,
  "avg_latency_ms":   42.0,
  "peak_vram_mb":     4096
}

Response 200

json
{ "ok": true }

GET /nodes/:id/hardware ​

Hardware profile reported by the node.

Response 200

json
{
  "gpu": "metal",
  "device_name": "Apple M4 Pro",
  "vram_mb": 32768
}

GET /nodes/:id/metrics ​

Per-node metrics timeseries (ring buffer, last ~60 samples at 10s intervals).

Response 200

json
{
  "node_id": "...",
  "samples": [
    { "tps": 45.2, "latency_ms": 22.1, "ts_ms": 1745123456000 }
  ]
}

DELETE /nodes/:id ​

Deregister a node (node-initiated clean exit).

Response 204 — no body


Metrics ​

GET /metrics/summary ​

Cluster-wide aggregate metrics.

Response 200

json
{
  "total_nodes": 2,
  "ready_nodes": 2,
  "total_tokens": 50000,
  "total_requests": 200,
  "avg_latency_ms": 38.5
}

GET /metrics/model/:model_id ​

Per-model metrics.

Response 200

json
{
  "model_id": "llama3.2:1b",
  "nodes":    2,
  "requests": 150,
  "tokens":   30000,
  "avg_latency_ms": 35.0
}

GET /metrics/history ​

Historical request rate (last N minutes, 1-minute buckets).

Response 200

json
{
  "buckets": [
    { "ts_ms": 1745123400000, "requests": 12, "tokens": 4800 }
  ]
}

Admin Endpoints ​

All admin routes require X-Admin-Token: <token>.

GET /admin/verify ​

Check token validity.

Response 200

json
{ "ok": true }

Response 403 if invalid.

POST /admin/nodes/:id/reassign ​

Change the layer range a node is responsible for.

Request body

json
{ "start_layer": 0, "end_layer": 8 }

Response 200

json
{ "ok": true, "node_id": "...", "start_layer": 0, "end_layer": 8 }

PATCH /admin/nodes/:id/state ​

Override node state.

Request body

json
{ "state": "ready" }

state values: "ready" · "idle"

Response 200

json
{ "ok": true, "node_id": "...", "state": "Ready" }

DELETE /admin/nodes/:id/force ​

Force-deregister a node (admin override — does not notify the node).

Response 200

json
{ "ok": true, "node_id": "..." }

Error Format ​

All errors return JSON:

json
{ "error": "human-readable description" }

Common status codes:

CodeMeaning
400Bad request / missing field
403Missing or invalid admin token
404Node or resource not found
409Conflict (e.g. duplicate node_id)
422Capability mismatch / unknown model
503No nodes available / layer gap

Free to run · Proprietary