Can I run Gemma 3 12B?
It is a dense model: all 12 billion parameters work on every token, so it has to fit in the card's memory to run at a usable speed. At Q4 the file is about 7 GB. Below is every graphics card we track, whether Gemma 3 12B runs on it with 32 GB of DDR4-3200 system RAM, and how fast — with a label on every number saying whether we measured it or computed it.
⭐ What it is good at, from our own use: general use, multilingual
On every card, with 32 GB of DDR4-3200 system RAM
| card | runs? | speed | how we know |
|---|---|---|---|
| No graphics card (CPU only) | runs | about 2 tok/s | ESTIMATED from memory bandwidth |
| GTX 1050 Ti | slow but works | 1–2 tok/s | ESTIMATED from memory bandwidth |
| GTX 970 | slow but works | about 2 tok/s | ESTIMATED from memory bandwidth |
| GTX 1060 6GB | slow but works | 2–3 tok/s | ESTIMATED from memory bandwidth |
| GTX 980 Ti | slow but works | 2–3 tok/s | ESTIMATED from memory bandwidth |
| RTX 2060 | slow but works | 2–3 tok/s | ESTIMATED from memory bandwidth |
| GTX 1660 Super | slow but works | 2–3 tok/s | ESTIMATED from memory bandwidth |
| RTX 3060 Laptop 6GB | slow but works | 2–3 tok/s | ESTIMATED from memory bandwidth |
| RTX 3050 8GB | slow but works | 3–4 tok/s | ESTIMATED from memory bandwidth |
| GTX 1070 | slow but works | 3–5 tok/s | ESTIMATED from memory bandwidth |
| RTX 4060 | slow but works | 3–5 tok/s | ESTIMATED from memory bandwidth |
| RTX 4060 Laptop 8GB | slow but works | 3–5 tok/s | ESTIMATED from memory bandwidth |
| GTX 1080 | slow but works | 4–5 tok/s | ESTIMATED from memory bandwidth |
| RTX 2070 | slow but works | 4–6 tok/s | ESTIMATED from memory bandwidth |
| RTX 2060 Super | slow but works | 4–6 tok/s | ESTIMATED from memory bandwidth |
| RTX 3070 | slow but works | 4–6 tok/s | ESTIMATED from memory bandwidth |
| RTX 3080 10GB | runs | 68–85 tok/s | ESTIMATED from memory bandwidth |
| GTX 1080 Ti | runs | 43–54 tok/s | ESTIMATED from memory bandwidth |
| RTX 2080 Ti | runs | 55–69 tok/s | ESTIMATED from memory bandwidth |
| RTX 3060 12GB | runs | 32–40 tok/s | ESTIMATED from memory bandwidth |
| RTX 4070 | runs | 45–56 tok/s | ESTIMATED from memory bandwidth |
| RTX 5070 12GB | runs | 60–75 tok/s | ESTIMATED from memory bandwidth |
| Tesla T4 16GB (server) | runs | 28–36 tok/s | ESTIMATED from memory bandwidth |
| RTX 4080 16GB | runs | 64–80 tok/s | ESTIMATED from memory bandwidth |
| RTX 5080 16GB | runs | 86–107 tok/s | ESTIMATED from memory bandwidth |
| Tesla M40 24GB (server card, very cheap) | runs | 26–32 tok/s | ESTIMATED from memory bandwidth |
| L4 24GB (server) | runs | 27–33 tok/s | ESTIMATED from memory bandwidth |
| Tesla P40 24GB (server card, cheap) | runs | 31–39 tok/s | ESTIMATED from memory bandwidth |
| RTX 3090 | runs | 83–104 tok/s | ESTIMATED from memory bandwidth |
| RTX 4090 | runs | 90–112 tok/s | ESTIMATED from memory bandwidth |
| Tesla V100 32GB (server) | runs | 80–100 tok/s | ESTIMATED from memory bandwidth |
| RTX 5090 32GB | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| A100 40GB (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| L40S 48GB (server) | runs | 77–96 tok/s | ESTIMATED from memory bandwidth |
| A100 80GB PCIe (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| A100 80GB SXM (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| H100 80GB SXM (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| H100 NVL 94GB (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| H200 141GB (server) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
Questions people actually ask
Can I run Gemma 3 12B on my computer?
With 32 GB of system RAM it runs on 39 of the 39 graphics cards we track — and even with no graphics card at all. The table below says which, how fast, and for the ones where it does not fit, why not. Every speed carries a label saying whether we measured it or computed it.
What is the smallest graphics card that runs Gemma 3 12B?
The card with the least VRAM on our list that still runs it is the GTX 1050 Ti (4 GB). Past that, for a model like this one, system RAM decides as much as the card does.
Does it need 16 GB or 32 GB of RAM?
With 16 GB of system RAM it runs on 39 cards; with 32 GB, on 39. The difference is what the extra RAM buys you for this model.
Where do these numbers come from?
From the same calculation our API answers with (`what_can_i_run` on our MCP endpoint), and a guard proves it says the same thing as the JavaScript on our calculator page. Rows labelled MEASURED ran on our own machines, with the machine and date written beside them; everything else says ESTIMATED and is given as a range.
Want it for your exact machine, with your memory speed and quantisation?
Ask the calculator, or call what_can_i_run on our
MCP endpoint — it is the same calculation.
Other models
- Llama 3.3 70B 70B
- Qwen3.6 35B-A3B 35B
- Qwen3 Coder 30B-A3B 30B
- Qwen 27B (dense) 27B
- Gemma 3 27B 27B
- Mistral Small 24B 24B
- GPT-OSS 20B 20B
- Phi-4 14B 14B
- Qwen 14B 14B
- Llama 3 8B 8B
- Mistral 7B 7B
- Qwen 4B (small brain) 4B
- Qwen 1.7B 1B