PDF to Markdown benchmark: pdfmd vs Marker vs Docling vs PyMuPDF
What 1,000 pages cost across the market, and what the money buys. Prices are each vendor's published rate, checked October 2, 2026. Quality and speed come from six public domain PDFs run through each engine on one machine and scored by a script, shown as percentages. The reference tables were checked cell by cell by an AI model, not yet by a person.
Value for money
The hosted services whose engines this benchmark measured, against pdfmd. Value for money is quality per dollar as a share of pdfmd's, where quality is the mean of table cells found and headings kept.
| Service | Per 1,000 pages | Table cells found | Headings kept | Speed vs pdfmd | Value for money vs pdfmd |
|---|---|---|---|---|---|
| pdfmd.dev | $0.89 | 95% | 98% | 100% | 100% |
| Datalab (hosted Marker), fast or balanced | $4.00 | 57% | 94% | 6% | 17% |
| IBM (hosted Docling), pay as you go | $4.00 | 86% | 99% | 3% | 21% |
Quality and speed are measured on each engine's open-source release at its default settings; a hosted service may run a newer model or different settings. Speed is pages per second as a share of pdfmd's on the same machine. The other services in the market table were not run, so they have a price and no score.
The market, per 1,000 pages
| Service | Mode | Output | Price | At volume | pdfmd costs |
|---|---|---|---|---|---|
| pdfmd.dev | Starter plan | Markdown | $0.89 | $0.75 Pro plan | n/a |
| Mathpix | Files API (batch) | Markdown | $1.50 | $1.00 over 30M pages a month | 41% less |
| Mistral | OCR 4.1, Batch API | Markdown | $2.00 | none published | 55% less |
| LlamaParse | Cost-effective | Markdown | $3.75 | none published | 76% less |
| AWS Textract | Layout | JSON | $4.00 | $3.00 over 1M pages a month | 78% less |
| Datalab (hosted Marker) | Fast or balanced | Markdown | $4.00 | $3.00 if you let them retain data | 78% less |
| IBM (hosted Docling) | Pay as you go | Markdown | $4.00 | none published | 78% less |
| Mistral | OCR 4.1 | Markdown | $4.00 | none published | 78% less |
| Mathpix | v3/pdf | Markdown | $5.00 | $3.50 over 1M pages a month | 82% less |
| Azure Document Intelligence | Layout | Markdown | $10.00 | $8.00 500K pages a month commitment | 91% less |
| Datalab (hosted Marker) | Accurate | Markdown | $10.00 | $7.50 if you let them retain data | 91% less |
| Google Document AI | Layout Parser | Layout blocks | $10.00 | $8.00 3-year savings plan | 91% less |
| Reducto | r-1 Parse | Markdown | $10.00 | $8.00 batch queue, 12-hour completion | 91% less |
| Upstage | Document Parse, Standard | Markdown or HTML | $10.00 | none published | 91% less |
| LlamaParse | Agentic | Markdown | $12.50 | none published | 93% less |
| AWS Textract | Tables and Layout | JSON | $15.00 | $10.00 over 1M pages a month | 94% less |
| Unstructured | All strategies | JSON elements | $15.00 | none published | 94% less |
Published pay-as-you-go prices in US dollars per 1,000 pages, read from each vendor's own pricing page or docs. US regions. A volume price needs the stated monthly volume, commitment or option. pdfmd is priced by monthly plan, so its rate per 1,000 pages assumes the plan's pages are used; the Starter rate is a launch price for a limited time. Left out: Text-only OCR (Azure Read, Google Enterprise Document OCR, AWS Textract DetectDocumentText, Upstage Document OCR) and LlamaParse Fast: they return text without tables or headings, so they are not comparable. LandingAI ADE: the price depends on how many characters come out, so there is no per-page price to compare. Adobe PDF Extract: no paid price is published.
Where pdfmd loses
- Speed, pages per second as a share of pdfmd's: Plain PyMuPDF text 3,076%, pdfmd 100%.
- Whole-text ordered recall of the table values: Docling 99.3%, pdfmd 97.9%.
- Whole-text ordered recall of the table values: Plain PyMuPDF text 100.0%, pdfmd 97.9%.
- Headings preserved: Docling 99.0%, pdfmd 98.1%.
- Ordered table-cell recall on Salient Statistics, United States (lithium): Marker 100.0%, pdfmd 95.8%.
Quality, every engine
| Engine | Ordered table-cell recall | Whole-text ordered recall | Headings kept | Speed vs pdfmd |
|---|---|---|---|---|
| pdfmd.dev (pymupdf4llm) | 95.2% (138/145) | 97.9% | 98.1% (101/103) | 100% |
| Marker | 57.2% (83/145) | 84.8% | 94.2% (97/103) | 6% |
| Docling | 86.2% (125/145) | 99.3% | 99.0% (102/103) | 3% |
| Plain PyMuPDF text | 0.0% (0/145) | 100.0% | 0.0% (0/103) | 3,076% |
Plain PyMuPDF writes text, not Markdown, so it has no tables or headings by construction: it is here as the floor, and its whole-text recall shows the text itself survives. Neither recall measure checks that a value sits in the right column; see Method.
Tables
| Table | Cells | pdfmd.dev (pymupdf4llm) | Marker | Docling | Plain PyMuPDF text |
|---|---|---|---|---|---|
| Salient Statistics, United States (lithium) page 1. Five year columns, withheld values written as W, a two line row label, and a > prefix on percentages. | 48 | 95.8% (46) | 100.0% (48) | 91.7% (44) | 0.0% (0) |
| Table 5: Page-level latencies for document indexing page 19. A spanning header over three columns and a bold, rule-separated header block in a LaTeX paper. | 25 | 92.0% (23) | 68.0% (17) | 80.0% (20) | 0.0% (0) |
| Table A-15. Alternative measures of labor underutilization page 27. Two level column header, nine data columns, and row labels wrapped over up to six lines with dot leaders. | 72 | 95.8% (69) | 25.0% (18) | 84.7% (61) | 0.0% (0) |
Headings
Scored on the three documents whose PDF outline lists real section headings: SEC Form 10-K, general instructions and form; ColPali: Efficient Document Retrieval with Vision Language Models (arXiv 2407.01449v6); NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0).
| Engine | sec-form-10k | colpali | nist-ai-rmf | Words run together | Missed, first few |
|---|---|---|---|---|---|
| pdfmd.dev (pymupdf4llm) | 39/41 | 32/32 | 30/30 | 0 | Item 1B. Unresolved Staff Comments.; Item 7. Management’s Discussion and Analysis of Financial Condition and Results of Operations. |
| Marker | 36/41 | 32/32 | 29/30 | 0 | C. Preparation of Report; D. Signature and Filing of Report; E. Disclosure With Respect to Foreign Subsidiaries |
| Docling | 40/41 | 32/32 | 30/30 | 14 | Part I |
| Plain PyMuPDF text | 0/41 | 0/32 | 0/30 | 0 | Form 10-K, ANNUAL Report Pursuant to Section 13 or 15(d) of the Securities Exchange Act of 1934; General Instructions; A. Rule as to Use of Form 10-K |
Method
- Ordered table-cell recall. Every table in an engine's output is a candidate, and so is every run of up to three adjacent tables, because some engines split one table at a header or page break. Ground truth rows are aligned to output rows in order, and cells within a row are matched in order. A value counts if it is inside a table, in the right row and in the right order. It is recall, not accuracy: empty cells are ignored on both sides, extra cells in the output cost nothing, and a value is not checked for the column it sits in. A row reading "WRONG | 2025 | 100" under "Year | Revenue" still scores 100 percent, because 2025 and 100 appear in order. Footnote markers, dot leaders, emphasis and dash styles are normalised away.
- Whole-text ordered recall. The same values looked for anywhere in the output text, in reading order, table or not. It shows whether the text survived at all, and is the only table measure plain text can score on.
- Headings kept. The PDF's own outline (its bookmarks) is the list of headings its authors declared. One is kept when a Markdown heading line matches it, keeping the order of both (the longest in-order match), ignoring section numbers, case and spacing. Spacing is ignored because the measure is whether the heading survived as a heading; an engine that glues a heading's words together ("TEXTUALRETRIEVALMETHODS") is counted as keeping it, and the number of such headings is shown in its own column as a text defect.
- Speed. Pages per second over the six documents, from the median of three timed runs per document, shown as a share of pdfmd's on the same machine. Models are loaded, and Marker's and Docling's first conversion done, before timing starts. The raw timings are in the committed results file.
- Price. Each vendor's published pay-as-you-go price per 1,000 pages, read from its own pricing page on October 2, 2026. The market median uses each vendor's cheapest mode that keeps tables or layout. pdfmd's price is read from the plans Stripe charges.
- Value for money. Quality (the mean of ordered table-cell recall and headings kept) divided by the price per 1,000 pages, as a share of pdfmd's. It is only shown where the benchmark measured the engine a hosted service runs.
- Choice of tables. The three tables were picked for structure that is hard to convert (a wrapped row label, a spanning header, a two level header with long wrapped labels) after looking at how pdfmd converted them, and before running the other engines.
Ground truth
Each table was transcribed from a 130 dpi render of its page (PyMuPDF get_pixmap) by reading the image, then cross-checked cell by cell against the PDF's own text layer (page.get_text()). Every value agreed on both reads. Claude (AI assistant) in the SEO sprint 1 session, 2026-09-24. This is a machine visual check, not a human one. Human sign-off: pending.
Documents
- Federal Reserve issues FOMC statement, 17 September 2025, Board of Governors of the Federal Reserve System. Board website policy: "Unless otherwise indicated, information on Board's website is in the public domain and may be copied and distributed without permission." Source. sha256 74f7ac6262051726
- USGS Mineral Commodity Summaries 2025: Lithium, U.S. Geological Survey. USGS policy: USGS-authored or produced data and information are considered to be in the U.S. public domain. Source. sha256 7ec5654133aa5ef6
- The Employment Situation, August 2025 (BLS news release, 5 September 2025), U.S. Bureau of Labor Statistics. BLS: "everything that we publish, both in hard copy and electronically, is in the public domain, except for previously copyrighted photographs and illustrations." Source. sha256 d91f1329ce9e7a13
- SEC Form 10-K, general instructions and form, U.S. Securities and Exchange Commission. A work of the U.S. Government (17 U.S.C. 105). SEC site policy: information on sec.gov "is considered public information and may be copied or further distributed by users of the web site without the SEC's permission." Source. sha256 65e8ce4f1364c735
- ColPali: Efficient Document Retrieval with Vision Language Models (arXiv 2407.01449v6), arXiv (authors: Faysse et al.). Released by its authors under CC0 1.0 Universal (public domain dedication), as stated on the arXiv abstract page. Source. sha256 852ec724f3d48316
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), National Institute of Standards and Technology. A work of the U.S. Government (17 U.S.C. 105). NIST: "information presented on NIST sites are considered public information and may be distributed or copied", except material marked as copyrighted. Source. sha256 7576edb531d98488
Engines and settings
- pdfmd.dev (pymupdf4llm): pymupdf4llm 1.28.2, pymupdf 1.28.2. code: workers/pdf/convert.py convert(); pymupdf4llm.to_markdown: {"page_chunks":true,"use_ocr":false,"force_ocr":false}; parser_version: 0.1.0
- Marker: marker-pdf 2.0.0, surya-ocr 0.22.1, transformers 5.17.0. converter: marker.converters.pdf.PdfConverter(artifact_dict=create_model_dict()); use_llm: false; output: markdown (default renderer); TORCH_DEVICE: unset (Marker picks the device); ocr_backend: Surya 2 through a local llama.cpp server: llama.cpp 0.5.0-dev (build 11169, commit 5cf3a3528); warm_up: one conversion of the FOMC statement inside load_s
- Docling: docling 2.130.0, docling-core 2.98.0, docling-ibm-models 4.0.3. converter: docling.document_converter.DocumentConverter() with default PdfPipelineOptions; do_ocr: true; ocr_options: OcrAutoOptions (kind=auto); the engine it picked is in ocr_modules_loaded; do_table_structure: true; table_structure_mode: TableFormerMode.ACCURATE; accelerator: num_threads=4 device='auto' cuda_use_flash_attention2=False; export: result.document.export_to_markdown(); warm_up: one conversion of the FOMC statement inside load_s
- Plain PyMuPDF text: pymupdf 1.28.2. call: page.get_text() per page, default flags
Machine
Apple M5 Pro, 18 cores, 48 GB memory, macOS 27.2 (26B5086k). Python 3.12.13. PyTorch engines had Apple Metal (MPS) available: true. The machine was not idle: other work was running during the timed runs, recorded below, so absolute times are slower than on a quiet machine. Every engine ran under the same conditions, one after another.
start 2026-09-24T22:48:01Z 17:48 up 2 days, 31 mins, 3 users, load averages: 3.65 4.44 5.45 %CPU COMM 81.8 opencode 24.8 Claude Helper (Renderer) 11.9 Claude Helper (Renderer) 11.8 Claude Helper 9.1 node end 2026-09-24T23:20:39Z 18:20 up 2 days, 1:03, 3 users, load averages: 5.34 5.76 5.75 %CPU COMM 37.0 opencode 8.8 node 5.9 ghostty 5.1 Claude Helper (Renderer) 3.3 Claude Helper
Reproduce
git checkout 6124feb33906 scripts/benchmark/setup.sh # one venv per engine, pinned versions scripts/benchmark/run_all.sh # fetch and verify PDFs, run engines one by one, score
Scored from commit 6124feb33906. Raw Markdown from every engine is committed under scripts/benchmark/outputs. See also the tables guide and the API docs.