Zember — a local AI sized to your graphics card
An AI that runs on your own computer. No account, no subscription, no internet after the first start. You pick the package by how much memory your graphics card has — the number is the spec.
Pick by your card
Every package holds the largest brain that fits entirely on that card. Nothing spills into system memory, so the speed on the box does not depend on what RAM you happen to own.
Zember Spark · 4 GB
Ministral 3 3B Q4_K_M · 2.0 GB · Apache 2.0
GTX 970, GTX 1050 Ti
contract test: 12 of 12 answers right
€39 · one key, for this package only
Buy Spark
31–54 tok/s
we ran this brain on an RTX 5070: 184 tok/s
ESTIMATED
BRAIN TESTED
Zember Glow · 6 GB
Qwen3.5 4B Q6_K · 3.3 GB · Apache 2.0
GTX 980 Ti, GTX 1060 6GB, RTX 2060, GTX 1660 Super, RTX 3060 Laptop 6GB
contract test: 12 of 12 answers right
€49 · one key, for this package only
Buy Glow
37–65 tok/s
we ran this brain on an RTX 5070: 129 tok/s
ESTIMATED
BRAIN TESTED
Zember Flame · 8 GB
Qwen3.5 9B Q4_K_M · 5.3 GB · Apache 2.0
GTX 1070, GTX 1080, RTX 2070, RTX 2060 Super, RTX 3070, RTX 3050 8GB, RTX 4060, RTX 4060 Laptop 8GB
contract test: 12 of 12 answers right
€59 · one key, for this package only
Buy Flame
32–65 tok/s
we ran this brain on an RTX 5070: 97 tok/s
ESTIMATED
BRAIN TESTED
Zember Torch · 10 GB
Ministral 3 14B Q4_K_M · 7.7 GB · Apache 2.0
RTX 3080 10GB, GTX 1080 Ti, RTX 2080 Ti
contract test: 12 of 12 answers right
€69 · one key, for this package only
Buy Torch
46–73 tok/s
we ran this brain on an RTX 5070: 64 tok/s
ESTIMATED
BRAIN TESTED
Zember Blaze · 12 GB
Phi-4 14B Q5_K_M · 9.7 GB · MIT
RTX 3060 12GB, RTX 4070, RTX 5070 12GB
contract test: 12 of 12 answers right
€79 · one key, for this package only
Buy Blaze
30–57 tok/s
we ran this brain on an RTX 5070: 57 tok/s
ESTIMATED
BRAIN TESTED
Zember Forge · 16 GB
Qwen3.8 27B UD-IQ4_XS · 13.3 GB · Apache 2.0
RTX 4080 16GB, RTX 5080 16GB
contract test: 12 of 12 answers right
€89 · one key, for this package only
Buy Forge
44–59 tok/s
we ran it on a 12 GB RTX 5070,
part in RAM: 7.0 tok/s
ESTIMATED
BRAIN TESTED
Zember Furnace · 24 GB
Granite 4.1 30B Q5_K_M · 19.1 GB · Apache 2.0
RTX 3090, RTX 4090, Tesla P40 24GB, Tesla M40 24GB
contract test: 12 of 12 answers right
40–43 tok/s
we ran it on a 12 GB RTX 5070,
part in RAM: 5.1 tok/s
ESTIMATED
BRAIN TESTED
Zember Crucible · 32 GB
Qwen3.6 35B-A3B Q6_K · 28.8 GB · Apache 2.0
RTX 5090 32GB
contract test: 12 of 12 answers right
25+ tok/s
we ran it on a 12 GB RTX 5070,
part in RAM: 24.9 tok/s
MEASURED
BRAIN TESTED
If your card is small but your RAM is not
These two are different: the brain is bigger than any card we list, so most of it sits in system memory. That means the speed is set by your RAM, not your card — a 6 GB card with fast memory beats a much more expensive card with slow memory. Nobody else prints this, because nobody else measured it.
Zember Inferno
Qwen3.6 35B-A3B Q4_K_M · 20.8 GB · spills into RAM
6 GB card → 32 GB RAM · 8 GB card → 24 GB RAM · 10 GB card → 24 GB RAM · 12 GB card → 24 GB RAM · 16 GB card → 16 GB RAM · 24 GB card → 8 GB RAM · 32 GB card → 8 GB RAM
estimated on other machines, tok/s: DDR4-2666: 13.7 · DDR4-3200: 17.0 · DDR5-4800: 25.7 · DDR5-6000: 32.0
contract test: 12 of 12 answers right
29.8 tok/s
measured: 12 GB RTX 5070
+ 64 GB RAM
MEASURED
Zember Wildfire
Qwen3.5 122B-A10B IQ3_S · 43.4 GB · spills into RAM
11 GB card → 48 GB RAM · 12 GB card → 48 GB RAM · 16 GB card → 48 GB RAM · 24 GB card → 32 GB RAM · 32 GB card → 24 GB RAM
estimated on other machines, tok/s: RTX 5090 + DDR5-6000: 34.0 · RTX 5090 + DDR4-3200: 19.0 · RTX 4090 + DDR5-6000: 21.7 · RTX 5070 + DDR5-6000: 14.6 · GTX 1080 Ti + DDR5-6000: 14.1
contract test: 12 of 12 answers right
13.6 tok/s
measured: 12 GB RTX 5070
+ 64 GB RAM
MEASURED
Why the numbers have labels
ESTIMATED is on every speed here, and it is honest: we do not own most of these cards. The figure comes from the card's memory bandwidth and an efficiency number we measured on five brains with the very engine in the box — not one borrowed from a chart.
BRAIN TESTED means we started that exact brain and made it write, on the card we own (an RTX 5070, 12 GB). For the small packages that card is faster than yours, so it is evidence of method, not a promise. For the big ones part of the brain sat in system RAM, so on the card they are made for they go faster. MEASURED is the speed we got on that machine.
Contract test: a 3-page contract built to trip it up (the real late-payment penalty in the last clause, wrong figures on the way, a later clause replacing an earlier one), four questions, each asked three times, through the very app in the box.
Every package carries a MEASURED.txt with its figures and how each one was produced. Each package is a separate product with its own key.