MarketAIVerse the universe of AI

Can I run Llama 3.3 70B?

It is a dense model: all 70 billion parameters work on every token, so it has to fit in the card's memory to run at a usable speed. At Q4 the file is about 42 GB. Below is every graphics card we track, whether Llama 3.3 70B runs on it with 32 GB of DDR4-3200 system RAM, and how fast — with a label on every number saying whether we measured it or computed it.

20 / 39cards run it with 32 GB of RAM
9 / 39with 16 GB
12 GBsmallest card that runs it

⭐ What it is good at, from our own use: high quality

On every card, with 32 GB of DDR4-3200 system RAM

cardruns?speedhow we know
No graphics card (CPU only)does not fit—The file is ~42.0 GB and you have 32 GB of RAM.
GTX 1050 Tidoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 970does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 1060 6GBdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 980 Tidoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 2060does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 1660 Superdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 3060 Laptop 6GBdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 3050 8GBdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 1070does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 4060does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 4060 Laptop 8GBdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 1080does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 2070does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 2060 Superdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 3070does not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 3080 10GBdoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
GTX 1080 Tidoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 2080 Tidoes not fit—The file is ~42.0 GB, and card plus RAM together do not hold it.
RTX 3060 12GBslow but worksunder 1 tok/sESTIMATED from memory bandwidth
RTX 4070slow but worksunder 1 tok/sESTIMATED from memory bandwidth
RTX 5070 12GBslow but worksunder 1 tok/sESTIMATED from memory bandwidth
Tesla T4 16GB (server)slow but worksunder 1 tok/sESTIMATED from memory bandwidth
RTX 4080 16GBslow but worksunder 1 tok/sESTIMATED from memory bandwidth
RTX 5080 16GBslow but worksunder 1 tok/sESTIMATED from memory bandwidth
Tesla M40 24GB (server card, very cheap)slow but worksunder 1 tok/sESTIMATED from memory bandwidth
L4 24GB (server)slow but worksunder 1 tok/sESTIMATED from memory bandwidth
Tesla P40 24GB (server card, cheap)slow but worksunder 1 tok/sESTIMATED from memory bandwidth
RTX 3090slow but works0–1 tok/sESTIMATED from memory bandwidth
RTX 4090slow but works0–1 tok/sESTIMATED from memory bandwidth
Tesla V100 32GB (server)slow but worksabout 1 tok/sESTIMATED from memory bandwidth
RTX 5090 32GBslow but worksabout 1 tok/sESTIMATED from memory bandwidth
A100 40GB (server)slow but worksabout 2 tok/sESTIMATED from memory bandwidth
L40S 48GB (server)runs14–18 tok/sESTIMATED from memory bandwidth
A100 80GB PCIe (server)runs32–41 tok/sESTIMATED from memory bandwidth
A100 80GB SXM (server)runs34–43 tok/sESTIMATED from memory bandwidth
H100 80GB SXM (server)runs56–70 tok/sESTIMATED from memory bandwidth
H100 NVL 94GB (server)runs65–82 tok/sESTIMATED from memory bandwidth
H200 141GB (server)runs80–101 tok/sESTIMATED from memory bandwidth

Questions people actually ask

Can I run Llama 3.3 70B on my computer?

With 32 GB of system RAM it runs on 20 of the 39 graphics cards we track. The table below says which, how fast, and for the ones where it does not fit, why not. Every speed carries a label saying whether we measured it or computed it.

What is the smallest graphics card that runs Llama 3.3 70B?

The card with the least VRAM on our list that still runs it is the RTX 3060 12GB (12 GB). Past that, for a model like this one, system RAM decides as much as the card does.

Does it need 16 GB or 32 GB of RAM?

With 16 GB of system RAM it runs on 9 cards; with 32 GB, on 20. The difference is what the extra RAM buys you for this model.

Where do these numbers come from?

From the same calculation our API answers with (`what_can_i_run` on our MCP endpoint), and a guard proves it says the same thing as the JavaScript on our calculator page. Rows labelled MEASURED ran on our own machines, with the machine and date written beside them; everything else says ESTIMATED and is given as a range.

Want it for your exact machine, with your memory speed and quantisation? Ask the calculator, or call what_can_i_run on our MCP endpoint — it is the same calculation.

Other models