Skip to content

Models ​

dppan runs GGUF models. The table below is the curated catalog — models validated end-to-end on the network. It is not a registry mirror: any GGUF from Ollama or Hugging Face can still be tried (dppan pull hf:<owner>/<repo>), the catalog just marks what we vouch for. Check it from the CLI anytime:

bash
dppan models          # table with local availability
dppan models --json   # for scripting
dppan models --tts    # speech models, catalogued separately
dppan models --music  # music models, catalogued separately

Looking for speech or music?

They are catalogued apart from the chat models below, because they are assembled from several files rather than sharded by layer. See Media: processing & generation for speech and Music for songs.

Supported models ​

0 / 0
ModelArchLayersSizeRegistriesManifest

The table is fetched live from dl.inclavate.io/models/catalog.json — the same catalog the CLI uses (dppan models caches it for 24 hours and works offline from a built-in copy). Models larger than one machine use sharded download.

Speech models are catalogued separately — see Media → Speech. They are not listed here because a speech model is a pair of files that always run together and never shards, so the layer count this table is organised around does not apply to one.

Music models are not here either, and for a stronger version of the same reason — see Media → Music. A music model is not one file or even a pair: it is four or five heterogeneous GGUFs pulled as a bundle into one directory, and which file plays which role is read from the files rather than declared. There is no single layer count to organise such a thing around, and no single quant either: quant is a per-component property, which is why dppan pull --component records one per file.

Coverage & roadmap ​

The engine supports 45 architecture families — enough to run 132 of the top-200 trending text-generation models on Hugging Face (audited 2026-09-02 by reading each model's GGUF architecture, not its name). Where the rest stands:

TierCoversStatus
Dense transformersLlama 3.x · Qwen 2.5/3 · Gemma 2/3/4 · Phi-3/4 · GLM-4 · SmolLM · Muse Glimmer 30B — plus their distills and finetunes✅ Supported
Dense classicsGPT-2 · Pythia · BLOOM · Falcon · Phi-2 · Nemotron · Command R7B✅ Supported
Gemma 4 E-seriesE2B · E4B — image and audio input; splits across at most two nodes✅ Supported
Mixture-of-ExpertsQwen3-MoE (A3B/235B) · gpt-oss · Gemma 4 A4B · OLMoE · Qwen1.5-MoE · GLM-4.5-Air and 355B-A32B · Mixtral 8x7B/8x22B · Llama 4 Scout & Maverick✅ Supported
MLA attentionDeepSeek-V2 · V2-Lite · Coder-V2-Lite · GigaChat3✅ Supported
State-space & hybridMamba (1 & 2) · FalconMamba · Codestral Mamba · Granite 4.0-H · Falcon-H1 (0.5B–34B)✅ Supported
Hybrid — single mixer per layerNemotron Nano 2 · Nemotron-H 47B Reasoning · Nemotron-3 MoE (30B-A3B – 550B-A55B) · LFM2 (350M–2.6B) · LFM2.5-2.6B · LFM2-8B-A1B · Jamba (900M · Mini 1.7 · Large 1.7)✅ Supported
Granite dense & MoEGranite 3.x/4.x dense (2B/8B) · Granite 4.2 (3B/8B/30B) · Granite-MoE a400m/a800m✅ Supported
Gated delta-netQwen3.5 (0.8B–27B) · Qwen3.6 · Qwen3.8 (27B + 2B/4B/9B distills) · Ornith 1.5 (9B · 35B-A3B) · Qwen3.5-MoE (A3B/A10B/A17B) · Qwen3-Next-80B · Qwen3-Coder-Next · Ling 3.0✅ Supported
MLA — frontierDeepSeek-V3 · R1 · V3.2 · Kimi K2 · Mistral-Large-3 · GLM-5/5.2 — 671B–1T📋 In Validation
Multimodal inputImage, audio and video frames on any model with a matching projector — Gemma 4 · Gemma 3 · Qwen2-VL · Qwen2.5-VL · Qwen3-VL · GLM-4.6V-Flash · DeepSeek-OCR · Muse Glimmer · Ultravox · Qwen3-ASR. Input only — the model answers in text✅ Supported
Speech to textOpenAI-compatible /v1/audio/transcriptions, plus a hands-free voice conversation in the dashboard. A chat model with an audio projector hears the clip directly; a text-only one gets a transcriber in front of it. See Media✅ Supported
Document inputPDF and text/Markdown files, optional OCR, native server-side video decode📋 Planned
Media generationSpeech and music ship. Speech via Qwen3-TTS (voice cloning, or nine named voices) and Orpheus/SNAC; music via MiniMax-Music3, ACE-Step 1.5 and YuE. See Media and the job queue. Image and video generation are not built🚧 In Progress

No dates are promised — items ship when they pass the same validation gate as everything in the catalog. Missing a model you need? Write to contact@inclavate.io — requests directly shape this list.

Model sources ​

dppan pull and dppan download resolve models from three places:

SourceRef formatAuth
Ollama registryllama3.3:70bnone
Hugging Face (GGUF repos)hf:bartowski/Llama-3.3-70B-Instruct-GGUF:Q6_Knone for public repos; your own HF_TOKEN for gated ones
Inclavate-hostedinclavate:<slug>[:<quant>]none
  • Hugging Face: only repos that ship GGUF files (bartowski, unsloth, lmstudio-community, …). The quant selector defaults to Q4_K_M, but any GGUF quantisation runs — pick another with hf:<owner>/<repo>:<QUANT> (e.g. :Q6_K, :Q8_0, :IQ4_XS); the engine keeps tensors in their native type and the right kernels are dispatched automatically. The quant shown in the catalog is simply the configuration we validated. Split quants (-00001-of-0000N.gguf) are merged automatically. For gated repos, accept the license on huggingface.co and set HF_TOKEN (or log in with huggingface-cli) — tokens are always your own.
  • Inclavate-hosted is for models with no community GGUF: converted and quantized once, centrally, then served from dl.inclavate.io — so nodes never have to quantize anything themselves.

Models published as a set of files ​

Most models are one GGUF. A few are not: ACE-Step 1.5 ships a diffusion transformer, a 5 Hz language model, a text encoder and an autoencoder in a single repository, and MiniMax-Music3 ships five components. Every file carries the same quant suffixes, so a plain dppan pull hf:…:Q8_0 cannot tell them apart — it refuses with the list of candidates rather than guessing at which of four models you meant.

--component names the one you want. ROLE is what the file is in your pipeline and is what gets recorded; SELECTOR finds it in the repository and defaults to ROLE, so you only write both when the publisher does not name its files after their job:

bash
# ACE-Step names its files after the checkpoint, so role and selector differ
dppan pull hf:Serveurperso/ACE-Step-1.5-GGUF:Q8_0 --component dit=acestep-v15-turbo
dppan pull hf:Serveurperso/ACE-Step-1.5-GGUF      --component vae

# MiniMax names them after the component, so the role alone is enough
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF:Q4_0  --component transformer
dppan pull hf:audio-cpp/MiniMax-Music3-GGUF       --component vocoder

One component per command. Each lands in ~/.dppan/models/bundles/<repo>/ alongside a bundle.dppan.json, which every pull extends — so a four-file model is assembled over four commands and the directory always says what it holds:

✓ component vocoder written: …/bundles/hf-audio-cpp-MiniMax-Music3-GGUF/vocoder.gguf
  bundle now holds 4: condition_encoder, rvq_depth_decoder (Q8_0), transformer (Q4_0), vocoder

Quant is per component, not per model — and it is read from the file that was selected, not from the one you asked for. That is not a nicety. These repositories genuinely mix precisions: MiniMax publishes its transformer at Q4_0 and its depth decoder at Q8_0, and ACE-Step publishes its autoencoder in BF16 only. Asking for :Q8_0 in the ACE-Step repo therefore gets you a Q8_0 diffusion transformer and a BF16 autoencoder, and the manifest records each as what it actually is rather than as what was requested. A component published in a single precision ignores the quant selector, because there is nothing to choose between — which is also why vocoder above needs no quant at all.

Some publishers export GGUFs that name their architecture in a way llama.cpp does not read. Where that is true and the fix is metadata rather than weights, the pull repairs the header on the way in and records that it did — the tensor payload is copied through byte for byte, so nothing is requantised.

--component works on Hugging Face repositories, requires the source to publish a sha256, and cannot be combined with --layers, --orch, --projector-only or --speech: a component is a whole file, not a layer slice. Point --out somewhere else to keep a bundle outside the model store.

Music models are assembled exactly this way, and they are served — see Music for a complete, copy-pasteable set of commands for one of them.

Integrity ​

A GGUF has no per-tensor checksums, so a partially-downloaded shard can't be checked against the source's whole-file hash. dppan's answer is a per-tensor manifest:

bash
dppan manifest llama3.3:70b -o 70b.manifest.json   # once, on a machine with the full file
dppan pull llama3.3:70b --layers 27-53 --verify --manifest 70b.manifest.json

With --verify, every downloaded tensor is hashed and compared; any mismatch names the tensor and deletes the shard. Without a manifest, shards are still structurally validated and fetched over TLS — fine for machines you control. Manifests published at dl.inclavate.io/manifests/<sha256>.json are picked up automatically, no --manifest flag needed.

Free to run · Proprietary