How we measure
This whole marketplace rests on one promise: every listing comes with a benchmark we ran. A promise without a procedure is just talk, so the procedure is written here, in the open, and anyone may dispute it.
A review says how good something is. A measurement says what happened, on what hardware, and when.
The seven rules
- 1. We run it, not the seller. If the seller runs it, it is not a benchmark, it is a screenshot.
- 2. Fixed inputs, the same for the whole category. And published. A benchmark on secret inputs cannot be checked by anyone.
- 3. Bad results get published too. A shelf where every benchmark is green is not an honest shelf.
- 4. Every number carries its label. Measured, computed, or nothing yet — and you can see which is which.
- 5. Measured = one number. Computed = a range. The shape of the number tells you what it is before you read it.
- 6. The machine and the date are written down. A number without a machine and a date is not published at all.
- 7. If the product changes, it gets re-measured. The old sheet stays in the history. Nothing is deleted.
And the eighth rule, which holds up all the others: the benchmark must be able to fail
Every input set contains at least one case we know in advance should fail — a corrupted file, an impossible request, an input outside the domain.
Why: a benchmark that comes back green every time does not prove the product is good. It only proves the benchmark cannot say „no". If even the case that was supposed to fail comes back green, the benchmark is broken, not the product brilliant — and then we publish nothing until the benchmark is fixed.
A guard you have only ever seen go green is not a guard.
The input set for document tools
Seven files, the same for any tool that takes a document in and puts a document out. MEASURED The figures on cases 6 and 7 are our own, measured on 13 September 2026 by opening the files with an ordinary library:
Download the set and check us
We are not asking you to take our word for it. These are exactly the files we run, on everyone. Take them, push them through your own tool, and compare with what we published. MEASURED The figures next to each file are our own, measured on 13 September 2026.
And why there are two cases, not one
We had designed a single corrupted file that „must be refused". Then we tested it, and nothing refused it. We broke the same PDF in six ways and opened it with an ordinary library. MEASURED 13 September 2026; every figure below is our own measurement:
So the right question is not „does it refuse?" but „does it notice content is missing?". A tool that only checks whether the file opened and how many pages it has reports success on a file with 0.4% left in it — and that reaches the customer.
⭐ The measurement changed our benchmark, not the other way round. Hence two cases: one silent, one clean.
For models and agents
We measure generation and prefill speed on a full context, not an empty one, and repeat after six hours of uptime.
And our own goods?
They go through exactly the same benchmark, and their sheets look the same. If we fail somewhere, it says so. Otherwise we would have no right to measure anyone else.