MarketAIVerse the universe of AI

PyMuPDF vs pdfplumber, measured

Both pull text out of a PDF. One is far faster, the other recovered a file the faster one gave up on, and the licence decides it for anyone building something to sell. These are our numbers, taken on Windows 11, 28 nuclee, doar procesor on 2026-09-16, against seven documents we publish so you can run it yourself and get your own.

39×faster, PyMuPDF, across the whole set
42×faster on the 60-page file alone
1file PyMuPDF refused and pdfplumber recovered

Every file, both tools

file, and what we required of itPyMuPDFpdfplumber charschars
1-simple.pdfmust come out whole0.0040 s0.1931 s10,57410,574
2-tables.pdftables must not be thrown away0.0043 s0.0956 s1,7631,763
3-text-in-images.pdflost without OCR - expected, not a fault0.0005 s0.0016 s00
4-diacritics.pdfevery diacritic must survive, cedilla forms included0.0027 s0.0312 s531494
5-large.pdfno sections lost0.0253 s1.0689 s65,50165,501
6-hollowed.pdfMUST notice the loss: almost no text, though it says 5 pages0.0015 s0.0010 s38—
7-no-header.pdfbroken at the start: refuse, or recover CORRECTLY - invented text would be the worst answer0.0008 s0.1776 s010,574

Marked rows are where the two disagree — different character counts or a different page count from the same file. Those rows are the whole point; the rest is just speed. MEASURED, Windows 11, 28 nuclee, doar procesor, 2026-09-16. Totals: 0.04 s against 1.57 s. Every file name above is a link — download them and get your own numbers.

The three answers, in the order that matters

Questions people actually ask

Which is faster, PyMuPDF or pdfplumber?

PyMuPDF, by a wide margin: 39 times on our seven-document set, and 42 times on the 60-page file alone. Measured on Windows 11, 28 nuclee, doar procesor, 2026-09-16.

So should I just use PyMuPDF?

Not necessarily, and this is the part a speed comparison hides. On the file with a damaged header PyMuPDF returned nothing at all, while pdfplumber recovered 10,574 characters - the same count it got from the intact version of that file. Refusing is a defensible answer; recovering the whole thing is a better one.

Can I use PyMuPDF in a product I sell?

Only if you are willing to publish your own source, or to buy a commercial licence. PyMuPDF is AGPL-3.0, which reaches even a product you only run as a web service. pdfplumber is MIT. This is the axis that decides it for most people, and it has nothing to do with speed.

What did you run this on?

Windows 11, 28 nuclee, doar procesor, on 2026-09-16, against a set of seven documents we publish so you can repeat it. Two of the seven are deliberately broken - a PDF with its text emptied and one with a damaged header - because a tool that only handles clean files tells you nothing.

Why is one file zero characters for both?

Because the text in it is inside images, and neither of these does OCR. That is the expected answer, not a failure - it is why Tesseract sits on this shelf too.

Why we bothered: an agent asked our API to compare exactly these two, on 18 September, and we had no page to give it. The requests that fail are the best list of what to write next — it is not what we think is interesting, it is what someone asked for and did not find.