You downloaded the model.
Now what?
Tell us what is in your computer. We will tell you what you can run — and whether it will actually be useful, which is not the same question.
VRAM calculators answer one question: does it fit? And they answer it wrongly for MoE models, because they do not know the experts can spill into system memory. That is why they told you „not possible" and you gave up.
What is in your computer
What you can run
Below is the same answer, worked out in advance for six common machines, with 32 GB of RAM and the usual Q4 compression. It is here in plain HTML on purpose: a program reading this page should not have to run our JavaScript to learn anything. Use the controls above for your own machine.
No graphics card (CPU only) 13 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | yes | ~7.1 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~7.1 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes | ~0.8 tok/s | ESTIMATED |
| Gemma 3 27B | yes | ~0.8 tok/s | ESTIMATED |
| Mistral Small 24B | yes | ~0.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~5.9 tok/s | ESTIMATED |
| Phi-4 14B | yes | ~1.5 tok/s | ESTIMATED |
| Qwen 14B | yes | ~1.5 tok/s | ESTIMATED |
| Gemma 3 12B | yes | ~1.8 tok/s | ESTIMATED |
| Llama 3 8B | yes | ~2.7 tok/s | ESTIMATED |
| Mistral 7B | yes | ~3.0 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | ~5.3 tok/s | ESTIMATED |
| Qwen 1.7B | yes | ~12.5 tok/s | ESTIMATED |
GTX 1050 Ti 12 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | no | — | NOT POSSIBLE |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 0.6–0.8 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 0.6–0.8 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 0.7–0.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Qwen 14B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Gemma 3 12B | yes, slowly | 1.4–2.0 tok/s | ESTIMATED |
| Llama 3 8B | yes, slowly | 2.3–3.2 tok/s | ESTIMATED |
| Mistral 7B | yes, slowly | 2.7–3.8 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes, slowly | 5.8–8.3 tok/s | ESTIMATED |
| Qwen 1.7B | yes | 49.4–71.4 tok/s | ESTIMATED |
GTX 1080 Ti 13 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.0–1.5 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.0–1.5 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.3–1.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 25.9–37.5 tok/s | ESTIMATED |
| Qwen 14B | yes | 25.9–37.5 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 30.3–43.7 tok/s | ESTIMATED |
| Llama 3 8B | yes | 45.4–65.5 tok/s | ESTIMATED |
| Mistral 7B | yes | 51.9–74.9 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 90.8–120.0 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 3060 12GB 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.3–0.4 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.1–1.6 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.1–1.6 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.4–2.1 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 19.3–27.9 tok/s | ESTIMATED |
| Qwen 14B | yes | 19.3–27.9 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 22.5–32.5 tok/s | ESTIMATED |
| Llama 3 8B | yes | 33.8–48.8 tok/s | ESTIMATED |
| Mistral 7B | yes | 38.6–55.7 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 67.5–97.5 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 4070 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.3–0.4 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.5–2.2 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 27.0–39.0 tok/s | ESTIMATED |
| Qwen 14B | yes | 27.0–39.0 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 31.5–45.5 tok/s | ESTIMATED |
| Llama 3 8B | yes | 47.3–68.3 tok/s | ESTIMATED |
| Mistral 7B | yes | 54.0–78.0 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 94.5–120.0 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 4090 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.4–0.6 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | over 120 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | over 120 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes | 28.0–40.4 tok/s | ESTIMATED |
| Gemma 3 27B | yes | 28.0–40.4 tok/s | ESTIMATED |
| Mistral Small 24B | yes | 31.5–45.5 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | over 120 tok/s | ESTIMATED |
| Phi-4 14B | yes | 54.0–78.0 tok/s | ESTIMATED |
| Qwen 14B | yes | 54.0–78.0 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 63.0–91.0 tok/s | ESTIMATED |
| Llama 3 8B | yes | 94.5–120.0 tok/s | ESTIMATED |
| Mistral 7B | yes | 108.0–120.0 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | over 120 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
MEASURED means it happened, on a named machine, on a named date. ESTIMATED is computed from a measurement, and the formula is published. A model marked NOT POSSIBLE does not fit; one marked yes, slowly is split between the card and system memory, which works but reads everything again for every word.
The MEASURED and ESTIMATED labels are not decoration. Here is exactly how we measure — and why every input set contains one case that must fail.
Why trust us
Because the numbers labelled MEASURED are not derived from spec-sheet formulas — they happened, on a specific machine, on a specific date, and both are written next to them. Everything else honestly says ESTIMATED, with the formula in plain sight. And where we do not know, we put no number at all.
Since late August we have kept a 35-billion-parameter model alive, in production, on a graphics card from 2017 — which is how we know what happens after six hours, not just in the first minute.
And if the answer is yes?
You downloaded the model. Now what?
How to keep it running, not just how to start it. With the measurement tables from a month of real production.
We set it up on your machine
We pick the model, set the right flags, and leave it running.