MarketAIVerse the universe of AI

How we measure

This whole marketplace rests on one promise: every listing comes with a benchmark we ran. A promise without a procedure is just talk, so the procedure is written here, in the open, and anyone may dispute it.

A benchmark is not a review. It is a measurement.
A review says how good something is. A measurement says what happened, on what hardware, and when.

The seven rules

And the eighth rule, which holds up all the others: the benchmark must be able to fail

Every input set contains at least one case we know in advance should fail — a corrupted file, an impossible request, an input outside the domain.

Why: a benchmark that comes back green every time does not prove the product is good. It only proves the benchmark cannot say „no". If even the case that was supposed to fail comes back green, the benchmark is broken, not the product brilliant — and then we publish nothing until the benchmark is fixed.

A guard you have only ever seen go green is not a guard.

The input set for document tools

Seven files, the same for any tool that takes a document in and puts a document out. MEASURED  The figures on cases 6 and 7 are our own, measured on 13 September 2026 by opening the files with an ordinary library:

1 · Plain document
clean text, 5 pages
must come out perfect
2 · Tables across pages
a technical manual is half tables
tables intact
3 · Text inside images
diagrams with writing on them
diagram text handled
4 · Accents and special characters
sounds trivial until they fall out
no character lost
5 · Large document
60+ pages, 30+ figures
finishes, loses no sections
6 · Hollowed out
cut to 4%: it opens, reports „5 pages", and holds 0.4% of the text
MUST NOTICE THE LOSS
7 · No header
zero recoverable pages
MUST REFUSE

Download the set and check us

We are not asking you to take our word for it. These are exactly the files we run, on everyone. Take them, push them through your own tool, and compare with what we published. MEASURED  The figures next to each file are our own, measured on 13 September 2026.

5 pages, clean text
29 KB
must come out perfect
tables across pages
47 KB
tables intact
no text layer at all: it is all drawing
70 KB
text inside figures
including the cedilla trap
3 KB
no character lost
60 pages, 36 figures
267 KB
no sections lost
opens, reports „5 pages", holds 0.4% of the text
1 KB
MUST NOTICE
zero recoverable pages
29 KB
MUST REFUSE
And the script that builds them: fa_intrari.py. The files are never written or edited by hand — they are regenerated, so they are identical for everyone and every time. If you cannot rebuild the set yourself, it is not a public set.

And why there are two cases, not one

We had designed a single corrupted file that „must be refused". Then we tested it, and nothing refused it. We broke the same PDF in six ways and opened it with an ordinary library. MEASURED  13 September 2026; every figure below is our own measurement:

Cut to 55%
6330 / 6345 characters
opens
Cut to 30%
52% of the text
opens
Cut to 8%
9.6% of the text
opens
Cut to 4%
and still reports „5 pages"
39 characters
opens
Scrambled bytes
more text than the original, i.e. garbage
8460 characters
opens
Header destroyed
0 pages
refuses
A PDF reader almost never refuses. It repairs, and hands back something.
So the right question is not „does it refuse?" but „does it notice content is missing?". A tool that only checks whether the file opened and how many pages it has reports success on a file with 0.4% left in it — and that reaches the customer.
⭐ The measurement changed our benchmark, not the other way round. Hence two cases: one silent, one clean.
And the worst possible result is not an error. It is a bad result returned silently. A tool that throws an error is usable — you know when not to trust it. A tool that quietly returns something broken is a trap, and that goes on the sheet in capital letters.

For models and agents

We measure generation and prefill speed on a full context, not an empty one, and repeat after six hours of uptime.

Measuring on an empty context is the most common lie in this field, and we learned it the hard way. MEASURED one setting gave us +15% on an empty context and −40% on the real one. If someone sends us numbers taken on an empty context, we re-measure.

And our own goods?

They go through exactly the same benchmark, and their sheets look the same. If we fail somewhere, it says so. Otherwise we would have no right to measure anyone else.

I want to be measured See the sheets on the shelf