Market
What is on the shelf right now, with the price in plain sight and the benchmark next to it.
The shelf
Zember Spark — Ministral 3 3B
Zember Glow — Qwen3.5 4B
Zember Flame — Qwen3.5 9B
Zember Torch — Ministral 3 14B
Zember Blaze — Phi-4 14B
Zember Forge — Qwen3.8 27B
Zember Furnace — Granite 4.1 30B
Zember Crucible — Qwen3.6 35B-A3B
Zember Inferno — Qwen3.6 35B-A3B
Zember Wildfire — Qwen3.5 122B-A10B
Technical manuals, translated and typeset
Data extraction from PDF
Product descriptions and search copy
Illustrations, logos, posters
llama.cpp ↗
Ollama ↗
What we found
Ran it for months as our model runner. Excellent to start with; we moved off it for the big quantised model, where we needed control over expert offloading. (2026-06 to 2026-08)
vLLM ↗
GPT4All ↗
What we found
We ran it. It is a desktop app rather than a server, so it is the one thing on this shelf you can start without opening a terminal - which matters more than it sounds if you are handing a local model to someone who does not want to learn one. (2026)
LM Studio ↗
What we found
We ran it. A desktop app rather than a server, so it is one of the few things here you can start without a terminal. (2026)
n8n ↗
Flowise ↗
Dify ↗
LangChain ↗
LangGraph ↗
CrewAI ↗
Haystack ↗
Hugging Face ↗
What we found
We pull every model through it, every week. It does its one job without drama. (used continuously through 2026)
ComfyUI ↗
What we found
The engine behind every video number on this shelf. One thing worth knowing before you trust it: a preset in its graph can silently replace your prompt instead of enriching it. (used continuously, last measured run 2026-09-06)
Stable Diffusion ↗
What we found
We ran it. Our own image work ended up on FLUX and on the small distilled models instead, so we have no quality claim to make here beyond having used it. (2026)
Wan 2.2 ↗
What we found
We ran it. On natural footage the 480p output plus an upscale held up; on fantasy prompts we rejected it. The 14B NF4 build peaked at 11.6 GB on a 12 GB card, so it fits, but with nothing to spare. (2026-07-11)
LTX-Video ↗
HunyuanVideo ↗
What we found
Peaked at 12.3 GB on a 12 GB card and never finished a clean run here. That is not a verdict on quality: it simply did not fit. (2026-07-11)
CogVideoX ↗
What we found
We ran it. It did not become part of our video pipeline — LTX-Video is what we settled on — so we have no quality claim to make beyond that. (2026-08)
Mochi 1 ↗
What we found
We ran it as an image-to-video candidate. It did not go further here; LTX-Video is the engine our own work runs on. (2026-08)
Whisper ↗
What we found
We ran the large-v3 weights on Romanian speech and they came back correct: diacritics right, company names kept, no invented sentences. We ran them through faster-whisper rather than the reference implementation, so the speed on the neighbouring card is the one we can point at. (2026-09-14)
Chroma ↗
Qdrant ↗
Weaviate ↗
Milvus ↗
LanceDB ↗
pgvector ↗
Vespa ↗
Docling ↗
Marker ↗
What we found
We ran it on our own documents, converting PDFs to markdown. (2026)
MinerU ↗
What we found
We ran it on our own documents. It is the one we reach for when a PDF has layout worth keeping rather than just text worth pulling. (2026)
Unstructured ↗
Tesseract OCR ↗
What we found
7 of 7 values came out correct on a scanned PDF with no text layer. It is what makes a scan readable at all. (2026-09-16)
PaddleOCR ↗
Surya ↗
OCRmyPDF ↗
PyMuPDF ↗
pdfplumber ↗
Camelot ↗
MarkItDown ↗
LayoutParser ↗
docTR ↗
GOT-OCR2.0 ↗
NLLB-200 ↗
OPUS-MT ↗
Argos Translate ↗
MADLAD-400 ↗
LibreTranslate ↗
SeamlessM4T ↗
faster-whisper ↗
whisper.cpp ↗
What we found
We ran it. Same weights as the speech-to-text numbers on this shelf, a different runtime - ours were taken through faster-whisper. (2026)
WhisperX ↗
What we found
We ran it. What it adds over plain transcription is word-level timing, which is the part that matters if you are cutting video to speech. (2026)
Piper ↗
Kokoro TTS ↗
Chatterbox ↗
F5-TTS ↗
MusicGen ↗
What we found
We ran it. ACE-Step is what we settled on for music here, so this one did not go further - that is a choice we made, not a defect we found. (2026)
Stable Audio Open ↗
What we found
We ran it. Same answer as MusicGen: we settled on ACE-Step, so this did not become part of our pipeline. (2026)
Demucs ↗
pyannote.audio ↗
Silero VAD ↗
Qwen2.5-VL ↗
What we found
We ran it. Worth knowing before you build on it: in this family the non-commercial licence sits on the small size while the 7B and 72B are Apache-2.0 - so the one that fits a modest card is the one you cannot sell on. See the licence traps page. (2026)
InternVL ↗
moondream ↗
Why we did not run it
Two versions, two answers. `moondream2` is clean; `moondream3-preview` forbids paid products that compete with it. Check which one you pulled. (2026-09-18)
Florence-2 ↗
Segment Anything 2 ↗
Ultralytics YOLO ↗
CLIP ↗
Grounding DINO ↗
rembg ↗
Real-ESRGAN ↗
BGE-M3 ↗
nomic-embed-text ↗
multilingual-e5 ↗
jina-embeddings-v3 ↗
GTE ↗
BGE Reranker ↗
Sentence Transformers ↗
ColBERT ↗
SPLADE ↗
Crawl4AI ↗
Firecrawl ↗
Playwright ↗
What we found
We ran it, and we still do: it drives a real browser for us, including the screen recordings we make of our own work. (2026)
browser-use ↗
What we found
We ran it - a browser an agent can drive by itself rather than by selectors written by hand. (2026)
Outlines ↗
Instructor ↗
Guidance ↗
Trafilatura ↗
Readability ↗
SearXNG ↗
ddgs ↗
E2B ↗
LiteLLM ↗
What we found
We ran it. One OpenAI-shaped API in front of many providers, which is the thing that lets you swap a model without touching the code that calls it. (2026)
FastMCP ↗
lm-evaluation-harness ↗
promptfoo ↗
Ragas ↗
DeepEval ↗
Langfuse ↗
Phoenix ↗
Guardrails ↗
garak ↗
Unsloth ↗
PEFT ↗
TRL ↗
LLaMA-Factory ↗
AutoAWQ ↗
GPTQModel ↗
What we found
We ran it, quantising models ourselves. Quantisation is the step that decides whether a model fits the card at all, and our own numbers on this shelf are almost all taken at q4 for exactly that reason. (2026)
LLM Compressor ↗
What we found
We ran it. Same job as the other quantiser here, with sparsity on top. (2026)
llama-swap ↗
What we found
We ran it. It swaps which model llama.cpp is holding without restarting the server - worth knowing, because a cold start of a large quantised model is the kind of wait we have measured in minutes, not seconds. (2026)
KoboldCpp ↗
Text Generation WebUI ↗
LocalAI ↗
SGLang ↗
Text Generation Inference ↗
MLX ↗
CTranslate2 ↗
Jan ↗
Open WebUI ↗
What we found
We ran it. A chat interface you host yourself in front of your own model, which is the piece that turns a local model into something you can hand to someone else without teaching them a command line. (2026)
LibreChat ↗
GPUStack ↗
RamaLama ↗
AutoGen ↗
LlamaIndex ↗
DSPy ↗
smolagents ↗
Pydantic AI ↗
OpenHands ↗
What we found
We ran it. An agent that writes and runs code on its own; what it costs is model calls, and that is the number to watch, not seconds. (2026)
Agno ↗
Langflow ↗
Activepieces ↗
Windmill ↗
Temporal ↗
FAISS ↗
USearch ↗
txtai ↗
Marqo ↗
Typesense ↗
Meilisearch ↗
OpenSearch ↗
What we found
We ran it, as a search and analytics store. (2026)
Mem0 ↗
GraphRAG ↗
LightRAG ↗
RAGFlow ↗
FLUX.1 ↗
What we found
We ran it. This is the image model our own work settled on — the numbers we published for image generation were taken on the schnell build at q4. (2026)
Stable Diffusion WebUI ↗
What we found
We ran it. Worth knowing before you build a product on it: it is AGPL, which is why the licence badge sits on this card. (2026)
SD WebUI Forge ↗
What we found
We ran it. Faster than the interface it forked from, and it carries the same AGPL terms — the badge on this card is not decoration. (2026)
InvokeAI ↗
Fooocus ↗
ControlNet ↗
Diffusers ↗
What we found
We ran it. It is the library underneath most of the image work here, rather than something you open. (2026)
SD.Next ↗
Kohya GUI ↗
ACE-Step ↗
yt-dlp ↗
What we found
Ran it as the first step of our speech-to-text chain, without a single failure. (2026-09-14)
AnimateDiff ↗
What we found
We ran it. Our own video work went to LTX-Video instead, so this did not become part of the pipeline. (2026)
FramePack ↗
FFmpeg ↗
What we found
Ran it to convert phone audio to 16 kHz mono WAV. Nothing surprising, which for this job is the highest praise. (2026-09-14)
MoviePy ↗
Auto-Editor ↗
RIFE ↗
Datasets ↗
DuckDB ↗
What we found
We ran it, as the analytical store for our own measurement data. (2026)
Polars ↗
DataTrove ↗
distilabel ↗
Argilla ↗
Label Studio ↗
What we found
We ran it, for labelling data by hand. (2026)
Great Expectations ↗
Presidio ↗
Rebuff ↗
ModelScan ↗
NeMo Guardrails ↗
Qwen3 ↗
What we found
We ran this family more than any other: the brain behind our own work has been a Qwen for months, and most measured numbers on this shelf were taken on one. (2026)
Llama ↗
What we found
We ran models from this family as candidate brains. (2026)
Mistral models ↗
What we found
We ran RoMistral-7B-Instruct (Q4_K_M) as a candidate brain. Its Romanian was genuinely good — the reason we kept it around for a while — and we later dropped it for a larger mixture-of-experts model, not because it was bad. (2026-06)
Gemma ↗
What we found
We ran models from this family as candidate brains. (2026)
OLMo ↗
What we found
We ran it. Fully open — weights, data and training code — which is rarer than the word 'open' usually means on a model page. (2026)
SmolLM ↗
What we found
We ran it. Small enough to run where nothing else fits, and we measured how fast that gets: our tiny-model numbers are on this shelf. (2026)
Hugging Face Hub CLI ↗
MCP reference servers ↗
MCP Python SDK ↗
MCP TypeScript SDK ↗
tiktoken ↗
Gradio ↗
Streamlit ↗
Chainlit ↗
Vercel AI SDK ↗
What we found
We ran it, building against it in TypeScript. (2026)
Transformers.js ↗
ONNX Runtime ↗
OpenVINO ↗
BentoML ↗
Transformers ↗
TripoSR ↗
InstantMesh ↗
TRELLIS ↗
Why we did not run it
Weighed on paper: architecture, size at our quantisation, and whether it fits the card we target. Never started. (2026-09-18)
Blender ↗
What we found
We ran it. It is the answer on this shelf to 3D and rigging, and it is on the shelf because an agent asked for exactly that and we had nothing. (2026)
Meshroom ↗
COLMAP ↗
Nerfstudio ↗
3D Gaussian Splatting ↗
OpenUSD ↗
Trimesh ↗
Depth Anything V2 ↗
MiDaS ↗
DUSt3R ↗
OpenPose ↗
MediaPipe ↗
Aider ↗
Cline ↗
Continue ↗
Tabby ↗
SWE-agent ↗
Tree-sitter ↗
Ruff ↗
Semgrep ↗
Chronos ↗
Lag-Llama ↗
StatsForecast ↗
Darts ↗
AutoGluon ↗
XGBoost ↗
LightGBM ↗
CatBoost ↗
TabPFN ↗
Neo4j ↗
Memgraph ↗
Kuzu ↗
cognee ↗
RealtimeSTT ↗
RealtimeTTS ↗
Moshi ↗
Parler-TTS ↗
ESPnet ↗
VideoLLaMA 3 ↗
InternVideo ↗
Decord ↗
PySceneDetect ↗
ESM ↗
Boltz ↗
RDKit ↗
DeepChem ↗
Docling Serve ↗
APScheduler ↗
Celery ↗
Redis ↗
MinIO ↗
Pandoc ↗
WeasyPrint ↗
Typst ↗
Playwright MCP ↗
Open Interpreter ↗
Semantic Kernel ↗
OpenAI Agents SDK ↗
MetaGPT ↗
CAMEL ↗
Letta ↗
Composio ↗
AgentScope ↗
Skyvern ↗
Stagehand ↗
AnythingLLM ↗
LobeHub ↗
Claude Code ↗
What we found
We ran it, and we still do: this shelf, this site and the tools behind it were built with a coding agent in a terminal. (2026)
Codex CLI ↗
What we found
We ran it. We keep it to reading rather than writing in our own work. (2026)
Goose ↗
Roo Code ↗
SWE-bench ↗
bolt.diy ↗
mypy ↗
Bandit ↗
TensorRT-LLM ↗
MLC LLM ↗
ExLlamaV2 ↗
llamafile ↗
KTransformers ↗
llama-cpp-python ↗
Xinference ↗
Sonar ↗
IPEX-LLM ↗
Triton Inference Server ↗
Ray ↗
OpenLLM ↗
Chatbox ↗
Text Embeddings Inference ↗
Infinity ↗
FastEmbed ↗
Model2Vec ↗
RAGatouille ↗
pgvecto.rs / VectorChord ↗
rerankers ↗
Verba ↗
kotaemon ↗
R2R ↗
EasyOCR ↗
RapidOCR ↗
olmOCR ↗
Zerox ↗
img2table ↗
pypdfium2 ↗
What we found
Ran it on our own seven-document test set. It did not make our published comparison, and we would rather say that than leave it out quietly. (2026-09-16)
pypdf ↗
Nougat ↗
Dolphin ↗
Detectron2 ↗
MMDetection ↗
Supervision ↗
OpenCLIP ↗
timm ↗
InsightFace ↗
DeepFace ↗
Kornia ↗
Albumentations ↗
BiRefNet ↗
DocLayout-YOLO ↗
SwarmUI ↗
IOPaint ↗
GFPGAN ↗
CodeFormer ↗
IP-Adapter ↗
InstantID ↗
PhotoMaker ↗
Krita AI Diffusion ↗
OneTrainer ↗
AI Toolkit ↗
What we found
We ran it, training our own image models. (2026)
sd-scripts ↗
What we found
We ran it, for training runs on image models. (2026)
Upscayl ↗
chaiNNer ↗
Open-Sora ↗
SadTalker ↗
Wav2Lip ↗
LivePortrait ↗
MuseTalk ↗
Video2X ↗
Deforum ↗
Coqui TTS ↗
Bark ↗
OpenVoice ↗
MeloTTS ↗
ChatTTS ↗
StyleTTS 2 ↗
GPT-SoVITS ↗
RVC ↗
NVIDIA NeMo ↗
What we found
We ran it, as a speech and language toolkit. (2026)
SpeechBrain ↗
Vosk ↗
Spleeter ↗
Audio Separator ↗
What we found
We ran it, splitting a track into stems. (2026)
librosa ↗
Descript Audio Codec ↗
Stable Audio Tools ↗
Opik ↗
TruLens ↗
Evidently ↗
Giskard ↗
AgentOps ↗
OpenLLMetry ↗
Inspect AI ↗
LLM Guard ↗
PyRIT ↗
Cleanlab ↗
Helicone ↗
MLflow ↗
DVC ↗
ZenML ↗
ClearML ↗
Metaflow ↗
Prefect ↗
Dagster ↗
Deep Lake ↗
FiftyOne ↗
CVAT ↗
doccano ↗
Megatron-LM ↗
DeepSpeed ↗
Accelerate ↗
torchtune ↗
verl ↗
OpenRLHF ↗
bitsandbytes ↗
LeRobot ↗
MuJoCo ↗
Genesis ↗
Gymnasium ↗
Stable-Baselines3 ↗
Isaac Lab ↗
Bullet / PyBullet ↗
InternLM ↗
Yi ↗
GLM-4 ↗
Kimi K2 ↗
What we found
We ran it. A large open model, and large is the operative word: it is well past what the cards we measure on can hold. (2026)
MiniCPM ↗
nanoGPT ↗
LLMs from scratch ↗
ScrapeGraphAI ↗
Jina Reader ↗
Markdownify ↗
newspaper4k ↗
LM Format Enforcer ↗
json_repair ↗
RouteLLM ↗
aisuite ↗
mcpo ↗
MCP Inspector ↗
Daytona ↗
What we found
We ran it, as a sandbox for code an agent writes. (2026)
Microsandbox ↗
FastAPI ↗
Pydantic ↗
How this market is laid out
These are not four markets. It is one picture with two axes: who sells and who buys. We open where we already have goods, and we say plainly where there is nothing yet.
We open with the two quadrants where we already have goods. The empty one stays in plain sight, honestly marked empty — a shelf that lies about being full is caught on the first click.