Every retrieval pipeline starts with the same unglamorous step: turning a pile of PDFs, scanned contracts, and slide decks into text a model can read. It is also where accuracy quietly leaks away, because a two-column paper, a merged-cell table, and a crooked fax scan each break a naive extractor in a different way.
Here is the number that reframes the choice this year. On OmniDocBench v1.6, a 1,651-page benchmark spanning 10 document types and 5 languages, the top scorer is PaddleOCR-VL-1.6 at 96.34 overall - an open-weights model with 0.9 billion parameters. Gemini 3 Pro scores 92.91. GPT-5.2 scores 86.59. A model small enough to run on a single GPU beats the frontier models at reading documents, and its weights are Apache 2.0.
That inverts the usual build-versus-buy math. "Just call an OCR API" was the sensible default for two years. In 2026 the best open models are more accurate and cheaper at volume, while the managed APIs are better at the one thing they were always good at: not making you run a GPU. The retrieval layer above this step is covered in our RAG platforms guide, and the store beneath it in best vector databases.
Why Parsing Decides RAG Quality
Parsing failures survive into production quietly, because nothing crashes.
Reading order is the first. Extract a two-column paper with a naive text layer and the columns interleave line by line, which reads as fluent nonsense to an embedding model. Tables are the second: flatten a financial statement into a paragraph and every number loses the row and column that gave it meaning, so the model confidently attributes Q3 revenue to Q1. Repeated headers and footers are the third, since a running title injected into every chunk adds noise to every embedding you store.
None of this shows up in an evaluation that only checks whether text was extracted. It shows up as an assistant that cites the right document and still gets the answer wrong.
Three Ways to Parse a Document
Pipeline libraries chain a layout model, an OCR engine, and an assembly step that emits Markdown or JSON. Docling and Marker work this way, run on CPU, and handle formats well beyond PDF.
Open-weight document VLMs replace that chain with one vision-language model that reads the page image and writes structured text directly. PaddleOCR-VL, MinerU, and DeepSeek-OCR 2 live here. This is where the accuracy is, and the cost is a GPU.
Managed APIs bill per page, trading cost and data residency for never thinking about model serving.
Open Source Parsers Worth Running
Docling

Docling is the safest default for a mixed corpus. IBM contributed it to the LF AI & Data Foundation in April 2025 and the codebase is MIT licensed, which removes the license question entirely. Version 2.123.0 landed August 26, 2026, at 65,600 GitHub stars.
Breadth is the real advantage. Docling ingests PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, LaTeX, email, and audio through ASR, then exports Markdown, HTML, or lossless JSON. It runs fully local, which is why it keeps appearing in air-gapped deployments. The project publishes about 1.5 pages per second on CPU alone, and the companion granite-docling-258M model (Apache 2.0) handles a page in one pass.
What it does well: format coverage, permissive licensing, genuine CPU-only operation.
What it does not do: top the leaderboards. Docling does not appear on OmniDocBench v1.6, and a dedicated VLM will beat it on difficult scans.
MinerU

MinerU from OpenDataLab is the highest-starred parser here at 78,600 stars, and 2026 brought the change that makes it viable commercially: on April 18, 2026 it moved off AGPL-3.0 to a custom Apache-2.0-based license. A separate commercial license is now required only above 100 million monthly active users or $20 million monthly revenue.
Version 3.4.5 shipped August 14, 2026. The backend spread is the selling point: the pure-CPU pipeline backend scores 86.47 on OmniDocBench v1.6 and needs 4GB, while hybrid-engine and vlm-engine reach about 95.3 with 8GB of VRAM. One tool covers the laptop prototype and the production run.
What it does well: near-leading accuracy, a real CPU fallback, a friendly license.
What it does not do: name things clearly. "MinerU 2.5" is the model, the tool is on 3.x.
PaddleOCR-VL

If you want the highest accuracy you can self-host, this is it. PaddleOCR-VL leads OmniDocBench v1.6 at 96.34 with a 0.9B Apache 2.0 model pairing a vision encoder with a small ERNIE language model. The parent project is the most-starred tool here at 88,300 stars, with v3.7.0 released in June 2026.
The catch is ecosystem buy-in: you adopt the PaddlePaddle framework rather than a plain PyTorch dependency, and much of the documentation skews toward Chinese-language sources. For a team dropping a parser into an existing Python service, that friction is real.
Marker

Marker from Datalab is the throughput option, and v2.0.0 on July 20, 2026 was a full rewrite around the Surya VLM that also added CPU support. On a single B200 it sustains 23.7 pages per second with OCR off, 7.4 in fast mode, and 2.9 in balanced mode, where it scores 76.0 on olmOCR-Bench.
Read the license first. The code is Apache 2.0, but the model weights use a modified AI Pubs Open RAIL-M license that excludes any organization with more than $5 million in prior-year revenue or $5 million raised in funding, and anyone building a competing product. Personal and research use are exempt.
The rest of the field
olmOCR 2 from Ai2 is fully Apache 2.0 including weights and scores 82.4 on its own olmOCR-Bench, but it is dormant with no commits since March 25, 2026. DeepSeek-OCR 2 (January 2026, Apache 2.0) is the interesting research direction, compressing text into vision tokens at roughly 10x with about 97% precision, a token-cost argument as much as an OCR one. Tesseract still ships at 5.5.3 and is fine for clean scans, but it has not kept pace.
Managed APIs and What They Cost

LlamaParse is the cheapest credible layout-aware option. Credits cost $1.25 per 1,000, and the four v2 tiers work out to $0.00125 per page for Fast, $0.00375 for Cost-effective, $0.0125 for Agentic, and $0.05625 for Agentic Plus. New accounts get 10,000 free credits, and re-parsing a file within 48 hours is free.

Reducto targets hard documents, and its pricing changes on September 1, 2026 from credits to per-product rates: $15 per 1,000 standard pages, $30 for complex pages, double for agentic modes, and $20 for Extract. The first 15,000 credits are free, and batching takes 20% off. It is the expensive end, for teams where a misread table has a dollar cost.
Mistral Document AI is the cheap, fast middle at $4 per 1,000 pages, halving to $2 through the Batch API. The current model is mistral-ocr-4-1, covering 170 languages. Accuracy depends on whose scorecard you read: Ai2's independently run olmOCR-Bench put the older Mistral OCR API at 72.0, while Mistral self-reports 85.20 for OCR 4, which would place it above every open model here. Weigh vendor-run numbers accordingly. Unusually for a hosted API, OCR 4 also ships as a single container for self-hosting, on enterprise plans only.
The hyperscalers converge tightly. Plain OCR is $1.50 per 1,000 pages on AWS Textract DetectDocumentText, Azure prebuilt-read, and Google Enterprise Document OCR alike, dropping to $0.60 at high volume. Layout-aware parsing jumps to $10 per 1,000 on both Azure prebuilt-layout and Google Layout Parser, while AWS charges $15 per 1,000 for tables. Google's free tier is the most usable at 1,000 pages per billing cycle. Unstructured sits at $0.015 per page with 10,000 free pages, and its open-source library stays Apache 2.0.
Comparison
| Tool | Model | License | OmniDocBench v1.6 | Cost | Best For |
|---|---|---|---|---|---|
| Docling | Self-host | MIT | Not listed | Free | Mixed formats, air-gapped |
| MinerU | Self-host | Apache-based custom | 86.47 CPU / 95.75 GPU | Free | Best all-round self-host |
| PaddleOCR-VL | Self-host | Apache-2.0 | 96.34 | Free | Maximum accuracy |
| Marker | Self-host | Apache-2.0 code, RAIL-M weights | 78.44 | Free under $5M | Throughput at scale |
| LlamaParse | API | Proprietary | Not listed | $0.00125-$0.056/page | Cheapest managed |
| Reducto | API | Proprietary | Not listed | $15-$30/1,000 pages | High-stakes documents |
| Mistral Document AI | API, self-host on enterprise | Proprietary | Not listed | $4/1,000 pages, $2 batch | Cheap bulk OCR |
| Cloud OCR (AWS/Azure/GCP) | API | Proprietary | Not listed | $1.50/1,000 plain | Already on that cloud |
How to Convert a PDF to RAG-Ready Markdown
Docling is the fastest path from PDF to chunkable text, and runs on CPU with no API key.
- Install it in a virtual environment, since Docling pulls in PyTorch.
- Convert the document. The first run downloads the layout and table models, so expect a wait.
- Inspect what you got before trusting it, especially the table count.
Bashpython3 -m venv .venv source .venv/bin/activate pip install docling
Pythonfrom pathlib import Path from docling.document_converter import DocumentConverter converter = DocumentConverter() doc = converter.convert("docling-report.pdf").document markdown = doc.export_to_markdown() Path("parsed.md").write_text(markdown, encoding="utf-8") print(f"pages: {len(doc.pages)}") print(f"tables: {len(doc.tables)}") print(f"chars: {len(markdown)}")
Running that against the 9-page Docling technical report on a laptop CPU prints:
pages: 9
tables: 3
chars: 34408
The table count is the number to watch. If a document you know contains tables reports zero, the parser dropped them into prose and your retrieval quality is already compromised.
How to Choose
- Can you run a GPU? If accuracy is the priority, PaddleOCR-VL. If you want one tool that degrades gracefully to CPU, MinerU.
- CPU only, or air-gapped? Docling, for the MIT license and format coverage.
- More than $5 million in revenue or funding? Marker's weights are off the table, however fast it is.
- Want zero infrastructure at low cost? LlamaParse Cost-effective at $0.00375 per page.
- Financially or legally consequential documents? Reducto, and budget for it.
Whatever you pick, evaluate on your own documents: a parser can be excellent on academic papers and poor on scanned invoices.
Related DevToolLab Tools
- Diff Checker - Diff two parsers' Markdown for the same page to see which one dropped a table or scrambled reading order.
- Regex Tester & Validator - Build the patterns that strip repeated headers, footers, and page numbers before you chunk.
- Markdown Editor - Preview parser output as rendered Markdown so broken tables are visible at a glance.
- Image to Base64 Converter - Encode a page image for a vision-model OCR endpoint when testing a VLM parser by hand.
Conclusion
The honest recommendation for most teams is MinerU if you can give it a GPU, Docling if you cannot. April's relicensing removed the one reason serious teams avoided MinerU, and its backend spread means the prototype and the production pipeline can be the same tool. Docling wins on license clarity and format breadth.
Reach for a managed API when your differentiator is elsewhere. LlamaParse at under half a cent per page costs less than the engineering time you would spend tuning a self-hosted stack, and Reducto earns its premium where an error is expensive. What genuinely changed this year is that self-hosting is no longer an accuracy compromise. A 0.9B open-weights model sitting above the frontier VLMs is a good problem for the rest of us to have.
Versions, benchmark scores, and pricing are as of August 27, 2026 and move quickly. Confirm on each project's repository or pricing page before you commit.
