What can I run on my machine?
Tell us what is in your computer. We will tell you what you can run — and whether it will actually be useful, which is not the same question.
VRAM calculators answer one question: does it fit? And they answer it wrongly for MoE models, because they do not know the experts can spill into system memory. That is why they told you „not possible" and you gave up.
What is in your computer
What you can run
Below is the same answer, worked out in advance for six common machines, with 32 GB of RAM and the usual Q4 compression. It is here in plain HTML on purpose: a program reading this page should not have to run our JavaScript to learn anything. Use the controls above for your own machine.
No graphics card (CPU only) 13 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | yes | ~7.1 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~7.1 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes | ~0.8 tok/s | ESTIMATED |
| Gemma 3 27B | yes | ~0.8 tok/s | ESTIMATED |
| Mistral Small 24B | yes | ~0.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~5.9 tok/s | ESTIMATED |
| Phi-4 14B | yes | ~1.5 tok/s | ESTIMATED |
| Qwen 14B | yes | ~1.5 tok/s | ESTIMATED |
| Gemma 3 12B | yes | ~1.8 tok/s | ESTIMATED |
| Llama 3 8B | yes | ~2.7 tok/s | ESTIMATED |
| Mistral 7B | yes | ~3.0 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | ~5.3 tok/s | ESTIMATED |
| Qwen 1.7B | yes | ~12.5 tok/s | ESTIMATED |
GTX 1050 Ti 12 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | no | — | NOT POSSIBLE |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 0.6–0.8 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 0.6–0.8 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 0.7–0.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Qwen 14B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Gemma 3 12B | yes, slowly | 1.4–2.0 tok/s | ESTIMATED |
| Llama 3 8B | yes, slowly | 2.3–3.2 tok/s | ESTIMATED |
| Mistral 7B | yes, slowly | 2.7–3.8 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes, slowly | 5.8–8.3 tok/s | ESTIMATED |
| Qwen 1.7B | yes | 51.7–64.6 tok/s | ESTIMATED |
GTX 1080 Ti 13 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | no | — | NOT POSSIBLE |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.0–1.5 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.0–1.5 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.3–1.9 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 36.6–45.7 tok/s | ESTIMATED |
| Qwen 14B | yes | 36.6–45.7 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 43.1–53.9 tok/s | ESTIMATED |
| Llama 3 8B | yes | 61.6–77.1 tok/s | ESTIMATED |
| Mistral 7B | yes | 66.7–83.4 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 97.5–120.0 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 3060 12GB 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.3–0.4 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.1–1.6 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.1–1.6 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.4–2.1 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 27.2–34.0 tok/s | ESTIMATED |
| Qwen 14B | yes | 27.2–34.0 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 32.1–40.1 tok/s | ESTIMATED |
| Llama 3 8B | yes | 45.9–57.3 tok/s | ESTIMATED |
| Mistral 7B | yes | 49.6–62.0 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 72.5–90.6 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 4070 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.3–0.4 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | ~17.0 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Gemma 3 27B | yes, slowly | 1.2–1.7 tok/s | ESTIMATED |
| Mistral Small 24B | yes, slowly | 1.5–2.2 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | ~14.2 tok/s | ESTIMATED |
| Phi-4 14B | yes | 38.1–47.6 tok/s | ESTIMATED |
| Qwen 14B | yes | 38.1–47.6 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 44.9–56.1 tok/s | ESTIMATED |
| Llama 3 8B | yes | 64.2–80.2 tok/s | ESTIMATED |
| Mistral 7B | yes | 69.5–86.8 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | 101.5–120.0 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
RTX 4090 14 of 14 models run
| Model | Runs? | Speed | Label |
|---|---|---|---|
| Llama 3.3 70B | yes, slowly | 0.4–0.6 tok/s | ESTIMATED |
| Qwen3.6 35B-A3B | yes | over 120 tok/s | ESTIMATED |
| Qwen3 Coder 30B-A3B | yes | over 120 tok/s | ESTIMATED |
| Qwen 27B (dense) | yes | 43.9–54.8 tok/s | ESTIMATED |
| Gemma 3 27B | yes | 43.9–54.8 tok/s | ESTIMATED |
| Mistral Small 24B | yes | 49.3–61.7 tok/s | ESTIMATED |
| GPT-OSS 20B | yes | over 120 tok/s | ESTIMATED |
| Phi-4 14B | yes | 76.2–95.3 tok/s | ESTIMATED |
| Qwen 14B | yes | 76.2–95.3 tok/s | ESTIMATED |
| Gemma 3 12B | yes | 89.8–112.3 tok/s | ESTIMATED |
| Llama 3 8B | yes | over 120 tok/s | ESTIMATED |
| Mistral 7B | yes | over 120 tok/s | ESTIMATED |
| Qwen 4B (small brain) | yes | over 120 tok/s | ESTIMATED |
| Qwen 1.7B | yes | over 120 tok/s | ESTIMATED |
MEASURED means it happened, on a named machine, on a named date. ESTIMATED is computed from a measurement, and the formula is published. A model marked NOT POSSIBLE does not fit; one marked yes, slowly is split between the card and system memory, which works but reads everything again for every word.
The MEASURED and ESTIMATED labels are not decoration. Here is exactly how we measure — and why every input set contains one case that must fail.
Every model, one page each
- Llama 3.3 70B 70B
- Qwen3.6 35B-A3B 35B
- Qwen3 Coder 30B-A3B 30B
- Qwen 27B (dense) 27B
- Gemma 3 27B 27B
- Mistral Small 24B 24B
- GPT-OSS 20B 20B
- Phi-4 14B 14B
- Qwen 14B 14B
- Gemma 3 12B 12B
- Llama 3 8B 8B
- Mistral 7B 7B
- Qwen 4B (small brain) 4B
- Qwen 1.7B 1B
Every card, one page each
Graphics cards
- No graphics card (CPU only) CPU only
- GTX 1050 Ti 4 GB
- GTX 970 4 GB
- GTX 1060 6GB 6 GB
- GTX 1660 Super 6 GB
- GTX 980 Ti 6 GB
- RTX 2060 6 GB
- RTX 3060 Laptop 6GB 6 GB
- GTX 1070 8 GB
- GTX 1080 8 GB
- RTX 2060 Super 8 GB
- RTX 2070 8 GB
- RTX 3050 8GB 8 GB
- RTX 3070 8 GB
- RTX 4060 8 GB
- RTX 4060 Laptop 8GB 8 GB
- RTX 3080 10GB 10 GB
- GTX 1080 Ti 11 GB
- RTX 2080 Ti 11 GB
- RTX 3060 12GB 12 GB
- RTX 4070 12 GB
- RTX 5070 12GB 12 GB
- RTX 4080 16GB 16 GB
- RTX 5080 16GB 16 GB
- RTX 3090 24 GB
- RTX 4090 24 GB
- RTX 5090 32GB 32 GB
Datacentre cards
- Tesla T4 16GB (server) 16 GB
- L4 24GB (server) 24 GB
- Tesla M40 24GB (server card, very cheap) 24 GB
- Tesla P40 24GB (server card, cheap) 24 GB
- Tesla V100 32GB (server) 32 GB
- A100 40GB (server) 40 GB
- L40S 48GB (server) 48 GB
- A100 80GB PCIe (server) 80 GB
- A100 80GB SXM (server) 80 GB
- H100 80GB SXM (server) 80 GB
- H100 NVL 94GB (server) 94 GB
- H200 141GB (server) 141 GB
Why trust us
Because the numbers labelled MEASURED are not derived from spec-sheet formulas — they happened, on a specific machine, on a specific date, and both are written next to them. Everything else honestly says ESTIMATED, with the formula in plain sight. And where we do not know, we put no number at all.
Since late August we have kept a 35-billion-parameter model alive, in production, on a graphics card from 2017 — which is how we know what happens after six hours, not just in the first minute.