PDF to Markdown in Python
Two ways to do this from Python. Call the API with requests and get JSON back, or run pymupdf4llm yourself, the open source library the pdfmd.dev converter is built on.
The example is NIST's AI Risk Management Framework, a 48 page public domain report, converted through the API's converter. The code below reproduces it and splits the result by heading.
A real example
NIST AI 100-1 (AI RMF 1.0), from the start of the Executive Summary, exactly as returned:
NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), National Institute of Standards and Technology.
Public domain: A work of the U.S. Government (17 U.S.C. 105). NIST: "information presented on NIST sites are considered public information and may be distributed or copied", except material marked as copyrighted. Source.
48 pages, 14 tables found, 0 scanned, 109,376 characters of Markdown. Converted by the service's own converter (parser 0.1.0, pymupdf4llm 1.28.2) when this page was built, in 4.31 s on the build machine.
## **Executive Summary** Artificial intelligence (AI) technologies have significant potential to transform society and people’s lives – from commerce and health to transportation and cybersecurity to the environment and our planet. AI technologies can drive inclusive economic growth and support scientific advancements that improve the conditions of our world. AI technologies, however, also pose risks that can negatively impact individuals, groups, organizations, communities, society, the environment, and the planet. Like risks for other types of technology, AI risks can emerge in a variety of ways and can be characterized as long- or short-term, highor low-probability, systemic or localized, and high- or low-impact. The AI RMF refers to an _AI system_ as an engineered or machine-based system that can, for a given set of objectives, generate outputs such as predictions, recommendations, or decisions influencing real or virtual environments. AI systems are designed to operate with varying levels of autonomy (Adapted from: OECD Recommendation on AI:2019; ISO/IEC 22989:2022). While there are myriad standards and best practices to help organizations mitigate the risks of traditional software or information-based systems, the risks posed by AI systems are in many ways unique (See Appendix B). AI systems, for example, may be trained on data that can change over time, sometimes significantly and unexpectedly, affecting system functionality and trustworthiness in ways that are hard to understand. AI systems and the contexts in which they are deployed are frequently complex, making it difficult to detect and respond to failures when they occur. AI systems are inherently socio-technical in nature, meaning they are influenced by societal dynamics and human behavior. AI risks – and benefits – can emerge from the interplay of technical aspects combined with societal factors related to how a system is used, its interactions with other AI systems, who operates it, and the social context in which it is deployed. These risks make AI a uniquely challenging technology to deploy and utilize both for organizations and within society. Without proper controls, AI systems can amplify, perpetuate, or exacerbate inequitable or undesirable outcomes for individuals and communities. With proper controls, AI systems can mitigate and manage inequitable outcomes.
Cut at a line boundary for length. Nothing between the first and last line shown has been edited.
This document on the benchmark
| Engine | Speed vs pdfmd | Headings kept |
|---|---|---|
| pdfmd.dev (pymupdf4llm) | 100% | 100.0% |
| Marker | 18% | 96.7% |
| Docling | 2% | 100.0% |
| Plain PyMuPDF text | 3,576% | 0.0% |
One machine, each engine's defaults. Speed is this document's conversion rate as a share of pdfmd's, from the median of three timed runs. Table-cell recall counts table values found in the right row and order; it does not check which column a value landed in. Method, all six documents and where we lose.
curl
curl -X POST https://pdfmd.dev/api/v1/convert \
-H "content-type: application/json" \
-d '{"url": "https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf"}'
# With an API key, add: -H "x-api-key: $PDFMD_API_KEY"Python
import requests
doc = requests.post(
"https://pdfmd.dev/api/v1/convert",
json={"url": "https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf"},
timeout=300,
).json()
md = doc["markdown"]
print(doc["title"], doc["pages"], "pages,", doc["tables"], "tables")
# Split by heading, keeping each heading with its text.
chunks, current = [], []
for line in md.splitlines():
if line.startswith("#") and current:
chunks.append("\n".join(current))
current = []
current.append(line)
chunks.append("\n".join(current))
print(len(chunks), "chunks")
# Or run the library the converter uses, locally (AGPL: check it suits your project):
# pip install pymupdf4llm
# import pymupdf4llm; md = pymupdf4llm.to_markdown("NIST.AI.100-1.pdf")Full reference: API docs. Limits and keys: pricing.
Limitations
- Long documents take longer, and the API answers only when the whole document is done. The benchmark below lists how long each document took on one machine.
- Running pymupdf4llm locally gives the same kind of Markdown but not necessarily byte for byte the same: the service pins one version and its own options.
- PyMuPDF and pymupdf4llm are AGPL licensed. Calling the API does not put that licence on your code; bundling the library can.
- Headings follow font sizes, so a report with many heading levels can come back flatter or deeper than its table of contents.
- The pageBreaks offsets in the response are currently two characters short for every page after the first, because they do not count the blank line that joins pages. Do not slice pages out of the Markdown with them until that is fixed; split by heading as above.