What can I run on a H200 141GB (server)?
This card has 141 GB of VRAM. Below is every language model we track, whether it runs on this card, and how fast — with a label on every number saying whether we measured it or computed it. Nothing here is a guess dressed up as a fact.
⭐ The difference between those two columns is RAM, not the card. For a mixture-of-experts model, past roughly 4 GB of VRAM the graphics card almost stops deciding anything — the experts stream from system memory on every token, so what matters is how much system RAM you have and how fast it is. Our calculation gives a GTX 1080 Ti and an RTX 5080 the same ESTIMATED 17 tok/s on the same model for exactly this reason — computed, not measured on a 5080.
The card itself, from the manufacturer's catalogue (not measured by us): 141 GB of VRAM, 4800 GB/s of memory bandwidth, released in 2024. The bandwidth is what the calculation uses for the part of a model that fits on the card.
With 16 GB of system RAM
| model | on this card | speed | how we know |
|---|---|---|---|
| Llama 3.3 70B | runs | 80–101 tok/s | ESTIMATED from memory bandwidth |
| Qwen3.6 35B-A3B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen3 Coder 30B-A3B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 27B (dense) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Gemma 3 27B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Mistral Small 24B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| GPT-OSS 20B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Phi-4 14B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 14B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Gemma 3 12B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Llama 3 8B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Mistral 7B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 4B (small brain) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 1.7B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
With 32 GB of system RAM
| model | on this card | speed | how we know |
|---|---|---|---|
| Llama 3.3 70B | runs | 80–101 tok/s | ESTIMATED from memory bandwidth |
| Qwen3.6 35B-A3B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen3 Coder 30B-A3B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 27B (dense) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Gemma 3 27B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Mistral Small 24B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| GPT-OSS 20B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Phi-4 14B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 14B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Gemma 3 12B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Llama 3 8B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Mistral 7B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 4B (small brain) | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
| Qwen 1.7B | runs | about 120 tok/s | ESTIMATED from memory bandwidth |
Questions people actually ask
What AI models can I run on a H200 141GB (server)?
With 32 GB of system RAM, 14 of the 14 models we track run on it. The table on this page says which, how fast, and for the ones that do not fit, why not. Every speed carries a label saying whether we measured it or computed it.
Is 141 GB of VRAM enough?
For a mixture-of-experts model, past roughly 4 GB of VRAM the card almost stops mattering - what decides is how much system RAM you have and how fast it is, because the experts stream from there on every token. Our calculation gives a GTX 1080 Ti and an RTX 5080 the same ESTIMATED 17 tok/s on the same model for exactly this reason. We have not run a 5080; the calculation is calibrated on the 1080 Ti we did run.
Why does an ordinary VRAM calculator say no?
Because it looks at the whole model file. In a mixture-of-experts model only the active experts need to be on the card at once, and the rest stream from system RAM. That is the single most common wrong 'no' people get.
Where do these numbers come from?
From the same calculation our API answers with, and a guard proves it says the same thing as the JavaScript on our own page - 882 comparisons, all identical. Where we actually ran the model ourselves the row says MEASURED with the machine and the date; everything else says ESTIMATED and is given as a range.
Have you actually run anything on a card like this?
No, and that matters. Our efficiency figure was calibrated on the cards we own and measured, up to 672 GB/s. This one is 7.1x faster than any of them, so every ESTIMATED number here is an extrapolation rather than something we watched happen. The arithmetic is the same and the shape should hold - but we would rather tell you than let a tidy number imply we were there.
Want it for your exact machine, with your memory speed? Ask the
calculator, or call what_can_i_run on our
MCP endpoint — it is the same calculation, and a guard proves
both say the same thing.
⭐ Or stop reading and try one. A table can tell you a small model answers at ten-odd tokens a second; it cannot tell you what that feels like, or how often that size is wrong. Run a 0.8B model here, free and without an account, on plain server CPU — the same kind of machine, without a graphics card at all.
Other cards
Graphics cards
- No graphics card (CPU only) CPU only
- GTX 1050 Ti 4 GB
- GTX 970 4 GB
- GTX 1060 6GB 6 GB
- GTX 1660 Super 6 GB
- GTX 980 Ti 6 GB
- RTX 2060 6 GB
- RTX 3060 Laptop 6GB 6 GB
- GTX 1070 8 GB
- GTX 1080 8 GB
- RTX 2060 Super 8 GB
- RTX 2070 8 GB
- RTX 3050 8GB 8 GB
- RTX 3070 8 GB
- RTX 4060 8 GB
- RTX 4060 Laptop 8GB 8 GB
- RTX 3080 10GB 10 GB
- GTX 1080 Ti 11 GB
- RTX 2080 Ti 11 GB
- RTX 3060 12GB 12 GB
- RTX 4070 12 GB
- RTX 5070 12GB 12 GB
- RTX 4080 16GB 16 GB
- RTX 5080 16GB 16 GB
- RTX 3090 24 GB
- RTX 4090 24 GB
- RTX 5090 32GB 32 GB
Datacentre cards
- Tesla T4 16GB (server) 16 GB
- L4 24GB (server) 24 GB
- Tesla M40 24GB (server card, very cheap) 24 GB
- Tesla P40 24GB (server card, cheap) 24 GB
- Tesla V100 32GB (server) 32 GB
- A100 40GB (server) 40 GB
- L40S 48GB (server) 48 GB
- A100 80GB PCIe (server) 80 GB
- A100 80GB SXM (server) 80 GB
- H100 80GB SXM (server) 80 GB
- H100 NVL 94GB (server) 94 GB