MarketAIVerse the universe of AI

Try a small language model

Below is a real model, running right now on the same ordinary server CPU that serves this page. No graphics card, no account, no key. Ask it something and watch how fast it answers - the speed you see is measured for your call, not an average we remembered.

Up to 600 characters. It writes back up to 160 tokens, then stops mid-sentence - that is the cap, not the model giving up. Up to 20 goes per hour from one address.
What you are looking at. Qwen3.5-0.8B at Q4_K_M, under llama.cpp, two CPU threads, 4096 token context. It is a weak model and it will be wrong - often, and confidently. That is the honest part: this page exists so you can see what this size really does, which includes watching it fail. Do not use anything it says as fact without checking it.

Questions people actually ask

Is this really free, and is there a catch?

It is free and there is no signup. The catch is the model: it is 0.8B parameters, which is tiny. It runs on a shared CPU with a cap of two cores, one caller at a time. If it is busy you are told how many seconds to wait rather than put in a queue.

What model is it, exactly?

Qwen3.5-0.8B quantised to Q4_K_M, served by llama.cpp on CPU only, two threads, 4096 token context. The size on disk is read off the file itself and shown in the box, so it cannot drift from what is actually loaded.

How fast is it really?

Measured on this machine at 14.0 tokens per second with two threads and no cap, and 9.3 to 12.6 tokens per second under the CPU cap it actually runs with. You do not have to take those numbers on faith: every answer reports the tokens per second measured for that exact call, straight from the engine.

Can I trust what it says?

No. A 0.8B model is wrong often, invents facts confidently, and loses track of long instructions. That is not a disclaimer we are obliged to write, it is the reason the box is here: most people have never watched a model this size fail, and reading an adjective for it is not the same as seeing it.

Why would anyone run a model this small?

Because it fits where a big one does not: it needs under a gigabyte of memory and no graphics card at all. For sorting text, pulling fields out of a form, or routing a request to the right place, small is often enough - and it is the difference between running on a laptop and renting a server.

Can my agent call this too, instead of using the page?

Yes. It is a tool called `try_a_model` on the MCP endpoint, and a plain `POST /api/v1/try` if you prefer HTTP. Same box, same limits, same measured numbers in the reply. Nothing here is for humans only.

Your agent can call it too

The same box is a tool on the machine side, because everything here has two faces over one set of data. Over MCP it is try_a_model; over plain HTTP it is this:

curl -s https://marketaiverse.com/api/v1/try \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "What is a small model good for?"}'

A GET on the same address describes the box without running anything, including whether it is up and what the limits are.