arXiv paper to Markdown
Pass an arXiv PDF URL and get the paper back as Markdown: section headings as # lines, tables as Markdown tables, running text in reading order across both columns.
arXiv papers carry their authors' own licences, and most are not public domain. The example here, the ColPali paper, is one its authors released under CC0, which is why it can be shown in full.
A real example
ColPali (arXiv 2407.01449v6, CC0 1.0), converted from its arXiv URL. The abstract, exactly as returned:
ColPali: Efficient Document Retrieval with Vision Language Models (arXiv 2407.01449v6), arXiv (authors: Faysse et al.).
Public domain: Released by its authors under CC0 1.0 Universal (public domain dedication), as stated on the arXiv abstract page. Source.
26 pages, 7 tables found, 0 scanned, 84,723 characters of Markdown. Converted by the service's own converter (parser 0.1.0, pymupdf4llm 1.28.2) when this page was built, in 2.91 s on the build machine.
## ABSTRACT Documents are visually rich structures that convey information through text, but also figures, page layouts, tables, or even fonts. Since modern retrieval systems mainly rely on the textual information they extract from document pages to index documents -often through lengthy and brittle processes-, they struggle to exploit key visual cues efficiently. This limits their capabilities in many practical document retrieval applications such as Retrieval Augmented Generation (RAG). To benchmark current systems on visually rich document retrieval, we introduce the Visual Document Retrieval Benchmark _ViDoRe_ , composed of various page-level retrieval tasks spanning multiple domains, languages, and practical settings. The inherent complexity and performance shortcomings of modern systems motivate a new concept; doing document retrieval by directly embedding the images of the document pages. We release _ColPali_ , a Vision Language Model trained to produce high-quality multi-vector embeddings from images of document pages. Combined with a late interaction matching mechanism, _ColPali_ largely outperforms modern document retrieval pipelines while being drastically simpler, faster and end-to-end trainable. We release models, data, code and benchmarks under open licenses at https://hf.co/vidore.
Nothing between the first and last line shown has been edited.
This document on the benchmark
| Engine | Speed vs pdfmd | Ordered table-cell recall | Headings kept |
|---|---|---|---|
| pdfmd.dev (pymupdf4llm) | 100% | 92.0% | 100.0% |
| Marker | 14% | 68.0% | 100.0% |
| Docling | 6% | 80.0% | 100.0% |
| Plain PyMuPDF text | 2,968% | 0.0% | 0.0% |
One machine, each engine's defaults. Speed is this document's conversion rate as a share of pdfmd's, from the median of three timed runs. Table-cell recall counts table values found in the right row and order; it does not check which column a value landed in. Method, all six documents and where we lose.
curl
curl -X POST https://pdfmd.dev/api/v1/convert \
-H "content-type: application/json" \
-d '{"url": "https://arxiv.org/pdf/2407.01449v6"}'
# With an API key, add: -H "x-api-key: $PDFMD_API_KEY"Python
import requests
def arxiv_to_markdown(arxiv_id: str) -> str:
resp = requests.post(
"https://pdfmd.dev/api/v1/convert",
json={"url": f"https://arxiv.org/pdf/{arxiv_id}"},
timeout=300,
)
resp.raise_for_status()
return resp.json()["markdown"]
md = arxiv_to_markdown("2407.01449v6")
# Section headings survive as "#" lines, so a paper splits on them.
sections = [s for s in md.split("\n## ") if s.strip()]
print(len(sections), "sections")Full reference: API docs. Limits and keys: pricing.
Limitations
- Equations come out as text, not LaTeX. Inline math and italic terms become underscore emphasis: the example has _ColPali_, and elsewhere in the paper a learning rate appears as 5 _e −_ 5.
- Figures are dropped; their captions stay as text.
- A table header set in spaced or rotated letters can be garbled. In this paper, Table 5's header "Indexing operation" came back as "Idi ti" and "nexng operaon".
- Use the arXiv PDF URL with a version (2407.01449v6), so the Markdown you store matches a fixed document.
- Check the paper's licence before republishing what you convert. Converting does not change it.