MarketAIVerse the universe of AI
guide · old hardware

You downloaded the model.
Now what?

Twenty gigabytes on disk, a card from 2017, and a VRAM calculator that told you it is not possible. It is. And the hard part is not starting it — it is keeping it running.

€25 Buy it — €25 First see what you can run
An old graphics card covered in dust, on a workbench
The hardware we are talking about. Not a data centre — the card in your drawer.

Why another guide

Guides exist. We read them. They all stop at the same place: „it loaded, it works". And almost all of them are written on 24 GB cards — that is, on hardware where anything runs anyway.

What they do not cover is what happens after six hours. When speed drops and you do not know why. When the cache is thrown away on every message. When a second person joins and everything halves. When the optimisation everybody recommends makes it worse.

The eleven chapters

8 pages, no filler. Every chapter carries at least one measured number, with the machine and the date next to it. MEASURED  Every figure in the list below is our own measurement, taken between 24 August and 12 September 2026 on the two machines described in the guide.

Their guides teach you to start it. This one teaches you to keep it running.

Proof that we know what we are talking about

Qwen3.6 35B-A3B on a GTX 1080 Ti (11 GB, from 2017)
7 hours straight, 500+ real answers: 23.86 → 23.73 tok/s
23.7 tok/s
MEASURED
Same model, no graphics card at all
And after 31.8 hours of uptime it had fallen to 2.05 — which is why the guide has a chapter on restarts
2.5 tok/s
MEASURED
These numbers are not derived from spec-sheet formulas. They happened, on our machines, on the dates written. Check what comes out for your hardware.

Read the first chapter, free

This is chapter 1 of 11, word for word as it is in the guide you would be buying — not a summary written for a sales page. If it is not worth €25 to you after reading it, it is not, and you keep your money.

1. The VRAM calculator lied to you

If you went to one of the seven VRAM calculators floating around and got „not possible",

remember one thing: they answer a single question, and they ask it wrong.

They compute: parameters × bits = how much memory you need. Does it fit?

For a 35-billion-parameter model at Q4 that is ~21 GB. Your card has 11. So: no.

And that is false, if the model is an MoE.

What an MoE is, in three lines

An ordinary („dense") model uses all its parameters for every token it writes. An MoE is split

into many „experts", and for each token only a few light up. In Qwen 35B-A3B, out of 35 billion

parameters, about 3 do the work on each token.

The part that always works — attention and the router, the thing that decides which expert fires —

is small. The huge part is the experts, and they are used rarely and one at a time.

So you do this

Put attention and the router on the card and leave the experts in system memory. For each

token, only the slice of expert that is needed gets read from RAM, not the whole model.

```

llama-server -m model.gguf --n-cpu-moe 999 -ngl 99 -np 1 -c 20480

```

--n-cpu-moe says how many experts stay on the CPU. Start with a large number to keep them all

there, then lower it until the card starts filling up.

Measured on the card, 30 August 2026: 23.7 tokens per second. On a card from 2017, with a

35-billion-parameter model that a VRAM calculator refuses you.

⭐ That is the whole move. The rest of this guide is about what happens after it starts.

---

How to get it

the guide

The guide and the files

The full text, the measurement tables and the ready-made commands.

€25Buy it
PDF, 8 pages, in English.
service

We set it up on your machine

We connect, choose, tune, and you end up with something that starts by itself.

€200COMING