PDF tables to Markdown for RAG, with a real example
Retrieval works best when a chunk carries its own structure: the heading it sits under, the table it belongs to, the page it came from. Plain text extraction throws that away. This page runs one real PDF through pdfmd.dev, shows what came out, and points at the parts you still need to check.
The example
A three page Federal Reserve H.15 statistical release (a public domain US government document, dated October 2016), converted from its URL:
curl -X POST https://pdfmd.dev/api/v1/convert \
-H "content-type: application/json" \
-d '{"url": "https://www.federalreserve.gov/releases/h15/current/h15.pdf"}'The response reported 3 pages, 0 scanned pages and 1 table, and took about 1.4 seconds. The release's prose came back as headings and paragraphs; the rates table came back as a GitHub flavoured Markdown table. The first rows:
|Itt|2016|2016|2016|2016|2016|Week|Ending|2016| |---|---|---|---|---|---|---|---|---| |nsrumens|Sep26|Sep27|Sep28|Sep29|Sep30|Sep30|Sep23|Sep| |Federal funds (effective)<sup>1 2 3</sup><br>|0.40|0.40|0.40|0.40|0.29|0.40|0.40|0.40| |Bank prime loan<sup>2 3 8</sup>|3.50|3.50|3.50|3.50|3.50|3.50|3.50|3.50| |<br>4-week|0.10|0.16|0.14|0.11|0.19|0.14|0.12|0.18| |3-month|0.25|0.26|0.27|0.26|0.28|0.26|0.24|0.29| |6-month|042|042|044|042|044|043|044|046|
What to check before you index it
The same excerpt shows the three defects worth testing for in your own documents:
- Split header text. The column heading “Instruments” came out as
Ittandnsrumensacross two rows, because the PDF draws it as rotated or spaced glyphs. - Lost decimal points. The 6-month row reads
042where the PDF shows 0.42. A number that looks plausible but is wrong is the most dangerous defect for retrieval, so spot check numeric cells against the source before trusting them. - Layout residue. Line breaks inside cells arrive as
<br>and footnote markers as<sup>. Strip or keep them deliberately; do not let them leak into embeddings by accident.
The title field also came from the PDF's own metadata (h15.dvi), not the visible heading. Use the first heading in the Markdown when the metadata title is unhelpful.
Chunking the output
- Split on Markdown headings first, so each chunk knows its section.
- Keep a table in one chunk with its header row. A table split mid-way loses its column meanings.
- Use
pageBreaksfrom the response (character offsets where each page starts) to attach a page number to every chunk, so an answer can cite where it came from. - Store the source URL and conversion date beside the chunks. Documents change; your index should say when it looked.
Limits
- Text PDFs only. Pages with almost no extractable text are counted in
scannedPagesand are not OCR'd. - Uploads up to 4 MB, a PDF at a URL up to 50 MB, and results up to 4 MB. Without a key the API allows 100 pages a day per address, up to 50 pages per file; a plan gives a monthly page allowance and files up to 400 pages. The converter on this site is counted separately, in files: 20 a day, up to 50 pages each and 200 pages in all.
- Table quality depends on how the PDF was drawn. Ruled, simple grids convert well; merged or rotated headers need checking, as above.
- No accuracy benchmark is claimed here. Measure on a sample of your own documents before indexing a collection.
Privacy
Our code holds an uploaded file and its Markdown in memory only and keeps neither, and there is no results URL. How long our hosting providers keep request logs we have not measured; see privacy and processing. For private documents, upload the file rather than hosting it at a public URL.