MarketAIVerse the universe of AI

What can I run on a Tesla V100 32GB (server)?

This card has 32 GB of VRAM. Below is every language model we track, whether it runs on this card, and how fast — with a label on every number saying whether we measured it or computed it. Nothing here is a guess dressed up as a fact.

14 / 14run with 32 GB of system RAM
14 / 14run with 16 GB
32 GBVRAM on this card

⭐ The difference between those two columns is RAM, not the card. For a mixture-of-experts model, past roughly 4 GB of VRAM the graphics card almost stops deciding anything — the experts stream from system memory on every token, so what matters is how much system RAM you have and how fast it is. Our calculation gives a GTX 1080 Ti and an RTX 5080 the same ESTIMATED 17 tok/s on the same model for exactly this reason — computed, not measured on a 5080.

The card itself, from the manufacturer's catalogue (not measured by us): 32 GB of VRAM, 900 GB/s of memory bandwidth, released in 2017. The bandwidth is what the calculation uses for the part of a model that fits on the card.

With 16 GB of system RAM

modelon this cardspeedhow we know
Llama 3.3 70Bslow but worksabout 1 tok/sESTIMATED from memory bandwidth
Qwen3.6 35B-A3Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen3 Coder 30B-A3Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 27B (dense)runs39–49 tok/sESTIMATED from memory bandwidth
Gemma 3 27Bruns39–49 tok/sESTIMATED from memory bandwidth
Mistral Small 24Bruns44–55 tok/sESTIMATED from memory bandwidth
GPT-OSS 20Brunsabout 120 tok/sESTIMATED from memory bandwidth
Phi-4 14Bruns68–85 tok/sESTIMATED from memory bandwidth
Qwen 14Bruns68–85 tok/sESTIMATED from memory bandwidth
Gemma 3 12Bruns80–100 tok/sESTIMATED from memory bandwidth
Llama 3 8Bruns115–120 tok/sESTIMATED from memory bandwidth
Mistral 7Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 4B (small brain)runsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 1.7Brunsabout 120 tok/sESTIMATED from memory bandwidth

With 32 GB of system RAM

modelon this cardspeedhow we know
Llama 3.3 70Bslow but worksabout 1 tok/sESTIMATED from memory bandwidth
Qwen3.6 35B-A3Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen3 Coder 30B-A3Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 27B (dense)runs39–49 tok/sESTIMATED from memory bandwidth
Gemma 3 27Bruns39–49 tok/sESTIMATED from memory bandwidth
Mistral Small 24Bruns44–55 tok/sESTIMATED from memory bandwidth
GPT-OSS 20Brunsabout 120 tok/sESTIMATED from memory bandwidth
Phi-4 14Bruns68–85 tok/sESTIMATED from memory bandwidth
Qwen 14Bruns68–85 tok/sESTIMATED from memory bandwidth
Gemma 3 12Bruns80–100 tok/sESTIMATED from memory bandwidth
Llama 3 8Bruns115–120 tok/sESTIMATED from memory bandwidth
Mistral 7Brunsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 4B (small brain)runsabout 120 tok/sESTIMATED from memory bandwidth
Qwen 1.7Brunsabout 120 tok/sESTIMATED from memory bandwidth

Questions people actually ask

What AI models can I run on a Tesla V100 32GB (server)?

With 32 GB of system RAM, 14 of the 14 models we track run on it. The table on this page says which, how fast, and for the ones that do not fit, why not. Every speed carries a label saying whether we measured it or computed it.

Is 32 GB of VRAM enough?

For a mixture-of-experts model, past roughly 4 GB of VRAM the card almost stops mattering - what decides is how much system RAM you have and how fast it is, because the experts stream from there on every token. Our calculation gives a GTX 1080 Ti and an RTX 5080 the same ESTIMATED 17 tok/s on the same model for exactly this reason. We have not run a 5080; the calculation is calibrated on the 1080 Ti we did run.

Why does an ordinary VRAM calculator say no?

Because it looks at the whole model file. In a mixture-of-experts model only the active experts need to be on the card at once, and the rest stream from system RAM. That is the single most common wrong 'no' people get.

Where do these numbers come from?

From the same calculation our API answers with, and a guard proves it says the same thing as the JavaScript on our own page - 882 comparisons, all identical. Where we actually ran the model ourselves the row says MEASURED with the machine and the date; everything else says ESTIMATED and is given as a range.

Want it for your exact machine, with your memory speed? Ask the calculator, or call what_can_i_run on our MCP endpoint — it is the same calculation, and a guard proves both say the same thing.

⭐ Or stop reading and try one. A table can tell you a small model answers at ten-odd tokens a second; it cannot tell you what that feels like, or how often that size is wrong. Run a 0.8B model here, free and without an account, on plain server CPU — the same kind of machine, without a graphics card at all.

Other cards

Graphics cards

Datacentre cards