PDF tables to Markdown
Tables come back as GitHub-flavoured Markdown tables, one row per line, which a language model or a Markdown renderer reads directly and which splits cleanly into chunks for retrieval.
The example is a table from a public domain USGS fact sheet, shown exactly as returned, defects included. Below it are table-cell recall numbers from our benchmark: how many values of three reference tables, checked by an AI model against the source and not yet by a person, each engine kept inside a table, in the right row and order. It does not check that each value sits in the right column.
A real example
The "Salient Statistics" table from the USGS Mineral Commodity Summaries 2025 lithium sheet, exactly as returned:
USGS Mineral Commodity Summaries 2025: Lithium, U.S. Geological Survey.
Public domain: USGS policy: USGS-authored or produced data and information are considered to be in the U.S. public domain. Source.
2 pages, 3 tables found, 0 scanned, 9,274 characters of Markdown. Converted by the service's own converter (parser 0.1.0, pymupdf4llm 1.28.2) when this page was built, in 0.31 s on the build machine.
|**Salient Statistics—United States:**|**2020**|**2021**|**2022**|**2023**|**2024**<sup>**e**</sup>| |---|---|---|---|---|---| |Production|W|W|W|W|W| |Imports for consumption|2,460|2,640|3,260|3,390|3,300| |Exports<br>|1,200|1,870|2,440|1,960|1,700| |Consumption, apparent<sup>1</sup>|W|W|W|W|W| |Price, annual average-real, battery-grade lithium carbonate,<br>|10,100|14,200|71,100|41,300|14,000| |dollars per metric ton<sup>2</sup>|||||| |Employment, mine and mill, number|70|70|70|70|70| |Net import reliance<sup>3</sup>as a percentage of apparent consumption|>50|>25|>25|>50|>50|
Nothing between the first and last line shown has been edited.
This document on the benchmark
| Engine | Speed vs pdfmd | Ordered table-cell recall |
|---|---|---|
| pdfmd.dev (pymupdf4llm) | 100% | 95.8% |
| Marker | 1% | 100.0% |
| Docling | 4% | 91.7% |
| Plain PyMuPDF text | 1,836% | 0.0% |
One machine, each engine's defaults. Speed is this document's conversion rate as a share of pdfmd's, from the median of three timed runs. Table-cell recall counts table values found in the right row and order; it does not check which column a value landed in. Method, all six documents and where we lose.
curl
curl -X POST https://pdfmd.dev/api/v1/convert \
-H "content-type: application/json" \
-d '{"url": "https://pubs.usgs.gov/periodicals/mcs2025/mcs2025-lithium.pdf"}'
# Pull just the tables out of the response with jq
curl -X POST https://pdfmd.dev/api/v1/convert \
-H "content-type: application/json" \
-d '{"url": "https://pubs.usgs.gov/periodicals/mcs2025/mcs2025-lithium.pdf"}' \
| jq -r '.markdown' | grep '^|'
# With an API key, add: -H "x-api-key: $PDFMD_API_KEY"Python
import re
import requests
doc = requests.post(
"https://pdfmd.dev/api/v1/convert",
json={"url": "https://pubs.usgs.gov/periodicals/mcs2025/mcs2025-lithium.pdf"},
timeout=180,
).json()
# Every Markdown table is a run of lines starting with "|".
tables = re.findall(r"(?:^\|.*\n?)+", doc["markdown"], flags=re.M)
for t in tables:
rows = [line.strip("|").split("|") for line in t.strip().splitlines()]
header, body = rows[0], [r for r in rows[2:]] # rows[1] is the --- separator
print(len(body), "rows:", header)Full reference: API docs. Limits and keys: pricing.
Limitations
- A row label that wraps onto a second line in the PDF becomes a second row. In the example, "dollars per metric ton" sits in a row of its own with empty cells, detached from its numbers.
- Footnote markers arrive as <sup> tags inside cells, and a <br> can appear where the PDF wrapped text inside a cell. Strip both before comparing values.
- Headers that span several columns are flattened, and a header set in spaced or rotated letters can come back garbled. The benchmark below shows where this costs table-cell recall.
- This Markdown endpoint does not read a table drawn as an image or a scanned page. The separate OCR endpoint reads English image text, but does not guarantee table structure.