Benchmark data
Document extraction benchmarks
We preserve the original 27-model benchmark as the baseline and add ten newer models to the same table. Compare quality, cost, speed, run date, and evidence status; document-processing engines remain separate.
LLM benchmark
Which model extracts best - and at what cost
March-August 2026 - 37 models, 4 document types, one table
Key findings
The original 27-model comparison remains intact as the baseline; ten newer routes extend it instead of replacing it.
Gemma 4 31B leads the latest audited cohort at 94.4% quality, while several historical models remain above that score.
Muse Glimmer 30B completed all four cases at 89.4%; Qwen3.6 27B reached 92.5% across three valid cases, with B4 left unscored.
Model quality, latency, and cost vary independently, so routing decisions should be based on representative documents rather than one aggregate rank.
CRM CSV
5 rows
Financial CSV
332 rows
Scanned invoice
OCR + tables
Investor table
32 rows
| Model / route | B1 | B2 | B3 | B4 | Valid | Quality | Valid time | Est. cost | Run | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4.1 minigpt-4.1-mini | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 994s | $0.090 | 2026-03-14 | Historical |
| Gemini 3.1 Flashgemini-3.1-flash | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 579s | $0.034 | 2026-03-14 | Historical |
| DeepSeek V3openrouter/deepseek-v3 | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 217s | ~$0.10 | 2026-03-14 | Historical |
| Gemma 4 31Bgoogle/gemma-4-31b-it | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 761.83s | $0.028998 | 2026-08-14 | Audited |
| Step 3.5 Flashopenrouter/step-3.5-flash | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 171s | Free | 2026-03-14 | Historical |
| DeepSeek V3 Freeopenrouter/deepseek-v3-free | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 182s | Free | 2026-03-14 | Historical |
| Kimi K2.5openrouter/kimi-k2.5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 212s | Not recorded | 2026-03-14 | Historical |
| MiMo V2 Flashopenrouter/mimo-v2-flash | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 221s | Not recorded | 2026-03-14 | Historical |
| GLM 5openrouter/glm-5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 246s | Not recorded | 2026-03-14 | Historical |
| Arcee Trinityopenrouter/arcee-trinity | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 259s | Free | 2026-03-14 | Historical |
| Kimi K2 (Groq)groq/kimi-k2 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 270s | $0.039 | 2026-03-14 | Historical |
| Qwen3 235Bopenrouter/qwen3-235b | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 325s | ~$0.08 | 2026-03-14 | Historical |
| Grok 4.1 Fastopenrouter/grok-4.1-fast | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 343s | Not recorded | 2026-03-14 | Historical |
| MiniMax M2.5openrouter/minimax-m2.5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 350s | Not recorded | 2026-03-14 | Historical |
| Claude Sonnet 4.6openrouter/claude-sonnet-4-6 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 560s | ~$1.10 | 2026-03-14 | Historical |
| GPT-5.4gpt-5.4 | 100% | 90% | 100% | 80% | 4 / 4 | 92.5% | 495s | $0.715 | 2026-03-14 | Historical |
| Qwen3.6 27Bqwen/qwen3.6-27b | 100% | 90% | 87.5% | Unscored | 3 / 4 | 92.5%* | 707.40s | $0.040607 | 2026-08-14 | Partial |
| GPT-5.6 Terraopenai/gpt-5.6-terra | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 102.19s | $0.226598 | 2026-08-14 | Audited |
| Muse Spark 1.2meta/muse-spark-1.2 | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 115.53s | $0.247949 | 2026-08-14 | Audited |
| GLM 5.2z-ai/glm-5.2 | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 207.88s | $0.135988 | 2026-08-14 | Audited |
| Gemini 3.7 Flashgoogle/gemini-3.7-flash | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 260.51s | $0.150934 | 2026-08-14 | Audited |
| Muse Glimmer 30Bmeta/muse-glimmer-30b | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 271.41s | $0.080912 | 2026-08-14 | Audited |
| Mistral Largeopenrouter/mistral-large | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 284s | ~$0.50 | 2026-03-14 | Historical |
| GPT-4.1 Nanogpt-4.1-nano | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 372s | $0.023 | 2026-03-14 | Historical |
| Gemini 3.6 Flashgoogle/gemini-3.6-flash | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 433.51s | $0.318894 | 2026-08-14 | Audited |
| Llama 4 Scout (Groq)groq/llama-4-scout | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 210s | $0.009 | 2026-03-14 | Historical |
| Llama 3.1 8B (Groq)groq/llama-3.1-8b | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 226s | Free tier | 2026-03-14 | Historical |
| Llama 3.3 70B (Groq)groq/llama-3.3-70b | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 289s | $0.022 | 2026-03-14 | Historical |
| Qwen3 32B (Groq)groq/qwen3-32b | 100% | 90% | 87.5% | 73% flaky | 3 / 4 stable | 87.6%* | 315s | Free tier | 2026-03-14 | Partial |
| GPT-4.1gpt-4.1 | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 348s | $0.458 | 2026-03-14 | Historical |
| GPT-5.3 Codexgpt-5.3-codex | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 462s | $0.721 | 2026-03-14 | Historical |
| GPT-5 Minigpt-5-mini | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 792s | $0.106 | 2026-03-14 | Historical |
| Claude Sonnet 5anthropic/claude-sonnet-5 | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 398.53s | $0.784292 | 2026-08-14 | Audited |
| Llama 4 Maverick (Groq)groq/llama-4-maverick | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 225s | Free tier | 2026-03-14 | Historical |
| GPT OSS 120B (Groq)groq/gpt-oss-120b | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 358s | $0.013 | 2026-03-14 | Historical |
| GPT OSS 20B (Groq)groq/gpt-oss-20b | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 407s | $0.014 | 2026-03-14 | Historical |
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash-0731 | 100% | Unscored | 87.5% | 60% | 3 / 4 | 82.5%* | 237.53s | $0.003618 | 2026-08-14 | Partial |
* Quality is averaged across valid cases. Unscored or flaky cases remain labeled and are not converted to zero.
Methodology
The original 27 models were tested on four document-extraction cases in March 2026. Ten models from the August 2026 cohort use the same B1-B4 columns and add strict runtime identity evidence. Run dates and evidence labels preserve the distinction between historical and audited results.
No single model is optimal for every document. Use the unified table as a shortlist, then benchmark candidate routes on your own document mix.
A note on comparability: All rows share the B1-B4 display schema. Historical and current cohorts retain their original run date and evidence status; document-processing engines are evaluated separately because they measure a different pipeline layer.
Engine benchmark
Document-processing engine comparison
April 2026 - 10 local engines measured, 19 engines rated
The engine layer sits below the LLM: it converts raw files to text before extraction begins. Engine choice affects OCR quality, table structure, multi-language coverage, and cost. This comparison covers 19 engines across local and cloud deployment.
Rating scale: Excellent / Good / Basic / Not supported
| Engine | Type | License | OCR | Tables | Multi-lang | Speed | Cost |
|---|---|---|---|---|---|---|---|
| Docling | Local | MIT | ★★☆ | ★★★ | ★★★ | Medium | Free |
| MinerU | Local | AGPL-3.0 | ★★★ | ★★★ | ★★☆ | Slow | Free |
| Chandra OCR 2 | Local / VLM | OpenRAIL-M | ★★★ | ★★★ | ★★★ | Slow | Free* |
| Marker | SaaS | Paid | ★★★ | ★★★ | ★★★ | Fast | ~$1/1K pages |
| PyMuPDF | Local | AGPL-3.0 | None | ★☆☆ | ★★★ | Ultra-fast | Free |
| PaddleOCR | Local | Apache-2.0 | ★★★ | ★★☆ | ★★★ | Medium | Free |
| Tesseract | Local | Apache-2.0 | ★★☆ | ★☆☆ | ★★★ | Medium | Free |
| EasyOCR | Local | Apache-2.0 | ★★★ | None | ★★★ | Medium | Free |
| Unstructured | Local / SaaS | Apache-2.0 | ★★☆ | ★★☆ | ★★☆ | Medium | Free / Paid API |
| LlamaParse | SaaS | Paid | ★★★ | ★★★ | ★★☆ | Fast | ~$3/1K pages |
| LiteParse | Local | Apache-2.0 | ★★☆ | ★★☆ | ★★☆ | Fast | Free |
| Mistral OCR | SaaS | Paid | ★★★ | ★★★ | ★★★ | Fast | ~$1/1K pages |
| Zerox | VLM | MIT | ★★★ | ★★☆ | ★★★ | Slow | VLM API cost |
| Nougat | Local | MIT | ★★☆ | ★★☆ | ★☆☆ | Slow | Free |
| Surya | Local | GPL-3.0 | ★★★ | ★★☆ | ★★★ | Medium | Free |
| AWS Textract | SaaS | Paid | ★★★ | ★★★ | ★★☆ | Fast | ~$1.50/1K pages |
| Google Document AI | SaaS | Paid | ★★★ | ★★★ | ★★★ | Fast | ~$1.50/1K pages |
| Azure Document Intelligence | SaaS | Paid | ★★★ | ★★★ | ★★★ | Fast | ~$1.50/1K pages |
| Firecrawl | SaaS | Paid | ★★☆ | ★★☆ | ★★★ | Fast | ~$1/1K pages |
Measured results (10 local engines, CPU)
Benchmarked on 2026-04-05 using synthetic PDF documents with known ground truth. All engines ran on CPU; no GPU. Times include model loading on first document.
| Engine | Avg time | Avg CER | Avg WER | Notes |
|---|---|---|---|---|
| PyMuPDF | 4 ms | 0.000 | 0.000 | Perfect on digital PDFs. No OCR capability. |
| Unstructured | 597 ms | 0.036 | 0.135 | Fast hybrid engine. Higher error on formatting. |
| Tesseract | 1,190 ms | 0.000 | 0.000 | Perfect on clean digital PDFs. 100+ languages. |
| PaddleOCR | 1,617 ms | 0.002 | 0.013 | Near-perfect accuracy. Best OCR at this speed. |
| Docling | 2,601 ms | 0.017 | 0.041 | Best accuracy-speed tradeoff among ML engines. |
| MinerU | 18,305 ms | 0.012 | 0.080 | Good CER among ML engines. Slow on CPU. |
| Surya | 32,959 ms | 0.014 | 0.027 | Lowest WER among ML engines. Slow on CPU. |
| Marker Local | 39,091 ms | 0.019 | 0.094 | Full document conversion with layout analysis. |
| EasyOCR | 30,311 ms | 0.000 | 0.000 | Perfect on single-page. Hangs on multi-page CPU. |
| Nougat | 127,380 ms | 1.000 | 1.000 | Designed for academic papers only. Wrong domain for simple text. |
CER = Character Error Rate. WER = Word Error Rate. Lower is better. 0.000 = perfect match.
Key takeaways
PyMuPDF is unmatched for digital PDFs: perfect accuracy at 4 ms average. Use it as the first-pass engine for text-layer documents.
PaddleOCR is the strongest OCR engine at this speed tier: 0.002 CER at 1.6 s on CPU.
Docling offers the best balance of speed and accuracy among ML-powered engines.
ML engines (Surya, Marker, MinerU) are designed for GPU and can be 5-20x faster with CUDA. CPU times shown here are worst-case.
Nougat is purpose-built for arXiv academic papers. It produces empty output on ordinary documents - not a defect, just wrong domain.
Bring your documents. We will find the right stack for them.
Book a 15-min discovery call. We will map extraction quality, cost, and data-residency requirements for your workload.