MarketAIVerse the universe of AI

Can I run Qwen3 Coder 30B-A3B?

It is a mixture-of-experts model: 30 billion parameters in total, but only about 3 billion work on each token. That is why it can run on cards far smaller than its file — the rest streams from system RAM, so RAM matters as much as the card. At Q4 the file is about 18 GB. Below is every graphics card we track, whether Qwen3 Coder 30B-A3B runs on it with 32 GB of DDR4-3200 system RAM, and how fast — with a label on every number saying whether we measured it or computed it.

39 / 39cards run it with 32 GB of RAM
14 / 39with 16 GB
4 GBsmallest card that runs it

⭐ What it is good at, from our own use: writing code

On every card, with 32 GB of DDR4-3200 system RAM

cardruns?speedhow we know
No graphics card (CPU only)runsabout 7 tok/sESTIMATED from memory bandwidth
GTX 1050 Tirunsabout 17 tok/sESTIMATED from memory bandwidth
GTX 970runsabout 17 tok/sESTIMATED from memory bandwidth
GTX 1060 6GBrunsabout 17 tok/sESTIMATED from memory bandwidth
GTX 980 Tirunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 2060runsabout 17 tok/sESTIMATED from memory bandwidth
GTX 1660 Superrunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 3060 Laptop 6GBrunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 3050 8GBrunsabout 17 tok/sESTIMATED from memory bandwidth
GTX 1070runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 4060runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 4060 Laptop 8GBrunsabout 17 tok/sESTIMATED from memory bandwidth
GTX 1080runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 2070runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 2060 Superrunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 3070runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 3080 10GBrunsabout 17 tok/sESTIMATED from memory bandwidth
GTX 1080 Tirunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 2080 Tirunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 3060 12GBrunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 4070runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 5070 12GBrunsabout 17 tok/sESTIMATED from memory bandwidth
Tesla T4 16GB (server)runsabout 17 tok/sESTIMATED from memory bandwidth
RTX 4080 16GBrunsabout 17 tok/sESTIMATED from memory bandwidth
RTX 5080 16GBrunsabout 17 tok/sESTIMATED from memory bandwidth
Tesla M40 24GB (server card, very cheap)runs72–104 tok/sESTIMATED from memory bandwidth
L4 24GB (server)runs75–108 tok/sESTIMATED from memory bandwidth
Tesla P40 24GB (server card, cheap)runs87–120 tok/sESTIMATED from memory bandwidth
RTX 3090runsabout 120 tok/sESTIMATED from memory bandwidth
RTX 4090runsabout 120 tok/sESTIMATED from memory bandwidth
Tesla V100 32GB (server)runsabout 120 tok/sESTIMATED from memory bandwidth
RTX 5090 32GBrunsabout 120 tok/sESTIMATED from memory bandwidth
A100 40GB (server)runsabout 120 tok/sESTIMATED from memory bandwidth
L40S 48GB (server)runsabout 120 tok/sESTIMATED from memory bandwidth
A100 80GB PCIe (server)runsabout 120 tok/sESTIMATED from memory bandwidth
A100 80GB SXM (server)runsabout 120 tok/sESTIMATED from memory bandwidth
H100 80GB SXM (server)runsabout 120 tok/sESTIMATED from memory bandwidth
H100 NVL 94GB (server)runsabout 120 tok/sESTIMATED from memory bandwidth
H200 141GB (server)runsabout 120 tok/sESTIMATED from memory bandwidth

Questions people actually ask

Can I run Qwen3 Coder 30B-A3B on my computer?

With 32 GB of system RAM it runs on 39 of the 39 graphics cards we track — and even with no graphics card at all. The table below says which, how fast, and for the ones where it does not fit, why not. Every speed carries a label saying whether we measured it or computed it.

What is the smallest graphics card that runs Qwen3 Coder 30B-A3B?

The card with the least VRAM on our list that still runs it is the GTX 1050 Ti (4 GB). Past that, for a model like this one, system RAM decides as much as the card does.

Does it need 16 GB or 32 GB of RAM?

With 16 GB of system RAM it runs on 14 cards; with 32 GB, on 39. The difference is what the extra RAM buys you for this model.

Where do these numbers come from?

From the same calculation our API answers with (`what_can_i_run` on our MCP endpoint), and a guard proves it says the same thing as the JavaScript on our calculator page. Rows labelled MEASURED ran on our own machines, with the machine and date written beside them; everything else says ESTIMATED and is given as a range.

Want it for your exact machine, with your memory speed and quantisation? Ask the calculator, or call what_can_i_run on our MCP endpoint — it is the same calculation.

Other models