A Dark Vector Cognition product
Guide

PDF to Markdown API

Send a PDF to POST /api/v1/convert, as a multipart file or as a URL, and the response is JSON: the Markdown, a title, the page count, how many pages look scanned, how many tables were found, and the character offset where each page starts.

Below is a real conversion of a public domain Federal Reserve statement through the same converter the API runs, then the calls to reproduce it and the things it does not do.

A real example

The FOMC statement of 17 September 2025, converted from its URL. The start of the Markdown, exactly as returned:

Federal Reserve issues FOMC statement, 17 September 2025, Board of Governors of the Federal Reserve System.

Public domain: Board website policy: "Unless otherwise indicated, information on Board's website is in the public domain and may be copied and distributed without permission." Source.

4 pages, 0 tables found, 0 scanned, 5,271 characters of Markdown. Converted by the service's own converter (parser 0.1.0, pymupdf4llm 1.28.2) when this page was built, in 0.22 s on the build machine.

For release at 2:00 p.m. EDT 

September 17, 2025 

Recent indicators suggest that growth of economic activity moderated in the first half of the year. Job gains have slowed, and the unemployment rate has edged up but remains low. Inflation has moved up and remains somewhat elevated. 

The Committee seeks to achieve maximum employment and inflation at the rate of 2 percent over the longer run. Uncertainty about the economic outlook remains elevated. The Committee is attentive to the risks to both sides of its dual mandate and judges that downside risks to employment have risen. 

In support of its goals and in light of the shift in the balance of risks, the Committee 

decided to lower the target range for the federal funds rate by 1/4 percentage point to 4 to 4-1/4 percent. In considering additional adjustments to the target range for the federal funds rate, 

the Committee will carefully assess incoming data, the evolving outlook, and the balance of risks. The Committee will continue reducing its holdings of Treasury securities and agency debt and agency mortgage‑backed securities. The Committee is strongly committed to supporting maximum employment and returning inflation to its 2 percent objective. 

In assessing the appropriate stance of monetary policy, the Committee will continue to monitor the implications of incoming information for the economic outlook. The Committee would be prepared to adjust the stance of monetary policy as appropriate if risks emerge that could impede the attainment of the Committee’s goals. The Committee’s assessments will take into account a wide range of information, including readings on labor market conditions, 

(more)

-2- 

inflation pressures and inflation expectations, and financial and international developments.

Nothing between the first and last line shown has been edited.

This document on the benchmark

EngineSpeed vs pdfmd
pdfmd.dev (pymupdf4llm)100%
Marker1%
Docling5%
Plain PyMuPDF text1,900%

One machine, each engine's defaults. Speed is this document's conversion rate as a share of pdfmd's, from the median of three timed runs. Table-cell recall counts table values found in the right row and order; it does not check which column a value landed in. Method, all six documents and where we lose.

curl

curl -X POST https://pdfmd.dev/api/v1/convert \
  -H "content-type: application/json" \
  -d '{"url": "https://www.federalreserve.gov/monetarypolicy/files/monetary20250917a1.pdf"}'

# Or upload a file you have
curl -X POST https://pdfmd.dev/api/v1/convert \
  -F "file=@statement.pdf"

# With an API key, add:  -H "x-api-key: $PDFMD_API_KEY"

Python

import requests

resp = requests.post(
    "https://pdfmd.dev/api/v1/convert",
    json={"url": "https://www.federalreserve.gov/monetarypolicy/files/monetary20250917a1.pdf"},
    headers={},  # {"x-api-key": "..."} if you have a key
    timeout=180,
)
if resp.status_code == 429:
    raise SystemExit("Rate limited; retry after " + resp.headers.get("X-RateLimit-Reset", "the reset time"))
resp.raise_for_status()
doc = resp.json()

print(doc["title"], doc["pages"], "pages,", doc["scannedPages"], "scanned")
markdown = doc["markdown"]

Full reference: API docs. Limits and keys: pricing.

Limitations

  • This endpoint reads text PDFs only. A page with almost no extractable text is counted in scannedPages and comes back nearly empty. For English scans, use the separate POST /api/v1/ocr endpoint.
  • Paragraphs can break where the PDF wraps a line. In the example, the sentence beginning "In support of its goals" is split in two at a line end.
  • One document per request, answered when the conversion finishes. There is no batch endpoint, queue or webhook.
  • There is no result URL to fetch later: the service does not keep the Markdown. Keep the JSON you get back.
  • pageBreaks is currently two characters short for every page after the first (it does not count the blank line that joins pages), so do not slice pages out of the Markdown with it until that is fixed.
  • Calls are rate limited per day. Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset, and GET /api/v1/me reports where you stand; the numbers are on the pricing page.

Other guides