MarketAIVerse the universe of AI

You downloaded the model.
Now what?

Tell us what is in your computer. We will tell you what you can run — and whether it will actually be useful, which is not the same question.

VRAM calculators answer one question: does it fit? And they answer it wrongly for MoE models, because they do not know the experts can spill into system memory. That is why they told you „not possible" and you gave up.

What is in your computer

What you can run

Below is the same answer, worked out in advance for six common machines, with 32 GB of RAM and the usual Q4 compression. It is here in plain HTML on purpose: a program reading this page should not have to run our JavaScript to learn anything. Use the controls above for your own machine.

No graphics card (CPU only) 13 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70BnoNOT POSSIBLE
Qwen3.6 35B-A3Byes~7.1 tok/sESTIMATED
Qwen3 Coder 30B-A3Byes~7.1 tok/sESTIMATED
Qwen 27B (dense)yes~0.8 tok/sESTIMATED
Gemma 3 27Byes~0.8 tok/sESTIMATED
Mistral Small 24Byes~0.9 tok/sESTIMATED
GPT-OSS 20Byes~5.9 tok/sESTIMATED
Phi-4 14Byes~1.5 tok/sESTIMATED
Qwen 14Byes~1.5 tok/sESTIMATED
Gemma 3 12Byes~1.8 tok/sESTIMATED
Llama 3 8Byes~2.7 tok/sESTIMATED
Mistral 7Byes~3.0 tok/sESTIMATED
Qwen 4B (small brain)yes~5.3 tok/sESTIMATED
Qwen 1.7Byes~12.5 tok/sESTIMATED

GTX 1050 Ti 12 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70BnoNOT POSSIBLE
Qwen3.6 35B-A3BnoNOT POSSIBLE
Qwen3 Coder 30B-A3Byes~17.0 tok/sESTIMATED
Qwen 27B (dense)yes, slowly0.6–0.8 tok/sESTIMATED
Gemma 3 27Byes, slowly0.6–0.8 tok/sESTIMATED
Mistral Small 24Byes, slowly0.7–0.9 tok/sESTIMATED
GPT-OSS 20Byes~14.2 tok/sESTIMATED
Phi-4 14Byes, slowly1.2–1.7 tok/sESTIMATED
Qwen 14Byes, slowly1.2–1.7 tok/sESTIMATED
Gemma 3 12Byes, slowly1.4–2.0 tok/sESTIMATED
Llama 3 8Byes, slowly2.3–3.2 tok/sESTIMATED
Mistral 7Byes, slowly2.7–3.8 tok/sESTIMATED
Qwen 4B (small brain)yes, slowly5.8–8.3 tok/sESTIMATED
Qwen 1.7Byes49.4–71.4 tok/sESTIMATED

GTX 1080 Ti 13 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70BnoNOT POSSIBLE
Qwen3.6 35B-A3Byes~17.0 tok/sESTIMATED
Qwen3 Coder 30B-A3Byes~17.0 tok/sESTIMATED
Qwen 27B (dense)yes, slowly1.0–1.5 tok/sESTIMATED
Gemma 3 27Byes, slowly1.0–1.5 tok/sESTIMATED
Mistral Small 24Byes, slowly1.3–1.9 tok/sESTIMATED
GPT-OSS 20Byes~14.2 tok/sESTIMATED
Phi-4 14Byes25.9–37.5 tok/sESTIMATED
Qwen 14Byes25.9–37.5 tok/sESTIMATED
Gemma 3 12Byes30.3–43.7 tok/sESTIMATED
Llama 3 8Byes45.4–65.5 tok/sESTIMATED
Mistral 7Byes51.9–74.9 tok/sESTIMATED
Qwen 4B (small brain)yes90.8–120.0 tok/sESTIMATED
Qwen 1.7Byesover 120 tok/sESTIMATED

RTX 3060 12GB 14 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70Byes, slowly0.3–0.4 tok/sESTIMATED
Qwen3.6 35B-A3Byes~17.0 tok/sESTIMATED
Qwen3 Coder 30B-A3Byes~17.0 tok/sESTIMATED
Qwen 27B (dense)yes, slowly1.1–1.6 tok/sESTIMATED
Gemma 3 27Byes, slowly1.1–1.6 tok/sESTIMATED
Mistral Small 24Byes, slowly1.4–2.1 tok/sESTIMATED
GPT-OSS 20Byes~14.2 tok/sESTIMATED
Phi-4 14Byes19.3–27.9 tok/sESTIMATED
Qwen 14Byes19.3–27.9 tok/sESTIMATED
Gemma 3 12Byes22.5–32.5 tok/sESTIMATED
Llama 3 8Byes33.8–48.8 tok/sESTIMATED
Mistral 7Byes38.6–55.7 tok/sESTIMATED
Qwen 4B (small brain)yes67.5–97.5 tok/sESTIMATED
Qwen 1.7Byesover 120 tok/sESTIMATED

RTX 4070 14 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70Byes, slowly0.3–0.4 tok/sESTIMATED
Qwen3.6 35B-A3Byes~17.0 tok/sESTIMATED
Qwen3 Coder 30B-A3Byes~17.0 tok/sESTIMATED
Qwen 27B (dense)yes, slowly1.2–1.7 tok/sESTIMATED
Gemma 3 27Byes, slowly1.2–1.7 tok/sESTIMATED
Mistral Small 24Byes, slowly1.5–2.2 tok/sESTIMATED
GPT-OSS 20Byes~14.2 tok/sESTIMATED
Phi-4 14Byes27.0–39.0 tok/sESTIMATED
Qwen 14Byes27.0–39.0 tok/sESTIMATED
Gemma 3 12Byes31.5–45.5 tok/sESTIMATED
Llama 3 8Byes47.3–68.3 tok/sESTIMATED
Mistral 7Byes54.0–78.0 tok/sESTIMATED
Qwen 4B (small brain)yes94.5–120.0 tok/sESTIMATED
Qwen 1.7Byesover 120 tok/sESTIMATED

RTX 4090 14 of 14 models run

ModelRuns?SpeedLabel
Llama 3.3 70Byes, slowly0.4–0.6 tok/sESTIMATED
Qwen3.6 35B-A3Byesover 120 tok/sESTIMATED
Qwen3 Coder 30B-A3Byesover 120 tok/sESTIMATED
Qwen 27B (dense)yes28.0–40.4 tok/sESTIMATED
Gemma 3 27Byes28.0–40.4 tok/sESTIMATED
Mistral Small 24Byes31.5–45.5 tok/sESTIMATED
GPT-OSS 20Byesover 120 tok/sESTIMATED
Phi-4 14Byes54.0–78.0 tok/sESTIMATED
Qwen 14Byes54.0–78.0 tok/sESTIMATED
Gemma 3 12Byes63.0–91.0 tok/sESTIMATED
Llama 3 8Byes94.5–120.0 tok/sESTIMATED
Mistral 7Byes108.0–120.0 tok/sESTIMATED
Qwen 4B (small brain)yesover 120 tok/sESTIMATED
Qwen 1.7Byesover 120 tok/sESTIMATED

MEASURED means it happened, on a named machine, on a named date. ESTIMATED is computed from a measurement, and the formula is published. A model marked NOT POSSIBLE does not fit; one marked yes, slowly is split between the card and system memory, which works but reads everything again for every word.

The MEASURED and ESTIMATED labels are not decoration. Here is exactly how we measure — and why every input set contains one case that must fail.

Why trust us

Because the numbers labelled MEASURED are not derived from spec-sheet formulas — they happened, on a specific machine, on a specific date, and both are written next to them. Everything else honestly says ESTIMATED, with the formula in plain sight. And where we do not know, we put no number at all.

Since late August we have kept a 35-billion-parameter model alive, in production, on a graphics card from 2017 — which is how we know what happens after six hours, not just in the first minute.

And if the answer is yes?

guide

You downloaded the model. Now what?

How to keep it running, not just how to start it. With the measurement tables from a month of real production.

service

We set it up on your machine

We pick the model, set the right flags, and leave it running.