PyMuPDF vs pdfplumber, measured
Both pull text out of a PDF. One is far faster, the other recovered a file the faster one gave up on, and the licence decides it for anyone building something to sell. These are our numbers, taken on Windows 11, 28 nuclee, doar procesor on 2026-09-16, against seven documents we publish so you can run it yourself and get your own.
Every file, both tools
| file, and what we required of it | PyMuPDF | pdfplumber | chars | chars |
|---|---|---|---|---|
1-simple.pdfmust come out whole | 0.0040 s | 0.1931 s | 10,574 | 10,574 |
2-tables.pdftables must not be thrown away | 0.0043 s | 0.0956 s | 1,763 | 1,763 |
3-text-in-images.pdflost without OCR - expected, not a fault | 0.0005 s | 0.0016 s | 0 | 0 |
4-diacritics.pdfevery diacritic must survive, cedilla forms included | 0.0027 s | 0.0312 s | 531 | 494 |
5-large.pdfno sections lost | 0.0253 s | 1.0689 s | 65,501 | 65,501 |
6-hollowed.pdfMUST notice the loss: almost no text, though it says 5 pages | 0.0015 s | 0.0010 s | 38 | — |
7-no-header.pdfbroken at the start: refuse, or recover CORRECTLY - invented text would be the worst answer | 0.0008 s | 0.1776 s | 0 | 10,574 |
Marked rows are where the two disagree — different character counts or a different page count from the same file. Those rows are the whole point; the rest is just speed. MEASURED, Windows 11, 28 nuclee, doar procesor, 2026-09-16. Totals: 0.04 s against 1.57 s. Every file name above is a link — download them and get your own numbers.
The three answers, in the order that matters
- Speed: PyMuPDF, and it is not close 39 times across the set. If you process documents in bulk this is the difference between a job that finishes and one you watch.
- Robustness: pdfplumber, on the one file that was broken Given a PDF with a damaged header, PyMuPDF returned nothing. pdfplumber returned 10,574 characters — the same count it got from the intact copy of that same file. Refusing is defensible. Recovering it is better.
- Licence: pdfplumber, and this one often ends the argument PyMuPDF is AGPL-3.0 — a product that includes it has to be open source too, even if you only run it as a web service. pdfplumber is MIT. More licence traps like this one.
Questions people actually ask
Which is faster, PyMuPDF or pdfplumber?
PyMuPDF, by a wide margin: 39 times on our seven-document set, and 42 times on the 60-page file alone. Measured on Windows 11, 28 nuclee, doar procesor, 2026-09-16.
So should I just use PyMuPDF?
Not necessarily, and this is the part a speed comparison hides. On the file with a damaged header PyMuPDF returned nothing at all, while pdfplumber recovered 10,574 characters - the same count it got from the intact version of that file. Refusing is a defensible answer; recovering the whole thing is a better one.
Can I use PyMuPDF in a product I sell?
Only if you are willing to publish your own source, or to buy a commercial licence. PyMuPDF is AGPL-3.0, which reaches even a product you only run as a web service. pdfplumber is MIT. This is the axis that decides it for most people, and it has nothing to do with speed.
What did you run this on?
Windows 11, 28 nuclee, doar procesor, on 2026-09-16, against a set of seven documents we publish so you can repeat it. Two of the seven are deliberately broken - a PDF with its text emptied and one with a damaged header - because a tool that only handles clean files tells you nothing.
Why is one file zero characters for both?
Because the text in it is inside images, and neither of these does OCR. That is the expected answer, not a failure - it is why Tesseract sits on this shelf too.
Why we bothered: an agent asked our API to compare exactly these two, on 18 September, and we had no page to give it. The requests that fail are the best list of what to write next — it is not what we think is interesting, it is what someone asked for and did not find.