A Dark Vector Cognition product
Guide

PDF tables to Markdown for RAG, with a real example

Retrieval works best when a chunk carries its own structure: the heading it sits under, the table it belongs to, the page it came from. Plain text extraction throws that away. This page runs one real PDF through pdfmd.dev, shows what came out, and points at the parts you still need to check.

The example

A three page Federal Reserve H.15 statistical release (a public domain US government document, dated October 2016), converted from its URL:

curl -X POST https://pdfmd.dev/api/v1/convert \
    -H "content-type: application/json" \
    -d '{"url": "https://www.federalreserve.gov/releases/h15/current/h15.pdf"}'

The response reported 3 pages, 0 scanned pages and 1 table, and took about 1.4 seconds. The release's prose came back as headings and paragraphs; the rates table came back as a GitHub flavoured Markdown table. The first rows:

|Itt|2016|2016|2016|2016|2016|Week|Ending|2016|
|---|---|---|---|---|---|---|---|---|
|nsrumens|Sep26|Sep27|Sep28|Sep29|Sep30|Sep30|Sep23|Sep|
|Federal funds (effective)<sup>1 2 3</sup><br>|0.40|0.40|0.40|0.40|0.29|0.40|0.40|0.40|
|Bank prime loan<sup>2 3 8</sup>|3.50|3.50|3.50|3.50|3.50|3.50|3.50|3.50|
|<br>4-week|0.10|0.16|0.14|0.11|0.19|0.14|0.12|0.18|
|3-month|0.25|0.26|0.27|0.26|0.28|0.26|0.24|0.29|
|6-month|042|042|044|042|044|043|044|046|

What to check before you index it

The same excerpt shows the three defects worth testing for in your own documents:

The title field also came from the PDF's own metadata (h15.dvi), not the visible heading. Use the first heading in the Markdown when the metadata title is unhelpful.

Chunking the output

Limits

Privacy

Our code holds an uploaded file and its Markdown in memory only and keeps neither, and there is no results URL. How long our hosting providers keep request logs we have not measured; see privacy and processing. For private documents, upload the file rather than hosting it at a public URL.

Convert a PDF now · API documentation