Pulling data from documents
Reading invoices, forms, contracts and statements and pulling out the pieces of information you need, such as totals, dates, names and each line of a table. A good result gets every field right. A missing value or a made-up value counts against the model.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- Reliability not yet measured
- 6 sources, 31 models with figures
- Sources last checked 2026-09-29
- Ring 4: no rerun evidence yet
- DB = DocuBench, EB = ExtractBench (LlamaIndex), PB = ParseBench, CX = ExtractBench (Contextual AI), IDP = IDP Core Bench, ODB = OmniDocBench
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| DocuBench | Claude Sonnet 5 Sent straight to the model | 91.7% | GPT-5.5 Sent straight to the model | 76.5% |
| ExtractBench (LlamaIndex) | GPT-6 Sol Coding agent (Codex), asked to show where each value came fromroughly tied, 3 setupsAlso roughly tied: GPT-5.5 Coding agent (Codex) 93.6% | 94.6% | GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from | 93.8% |
| ParseBench | Claude Opus 5.5 Sent straight to the model, effort highroughly tied | 79.85 points | Claude Fable 5.1 Sent straight to the model, default setting | 78.92 points |
| IDP Core Bench | Gemini 3 Pro (preview) Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: Claude Opus 4.6 Sent straight to the model, default settings 81.1% | 81.8% | Claude Sonnet 4.6 Sent straight to the model, default settings | 81.2% |
| OmniDocBench | Gemini 3 Pro (preview) Sent straight to the modelroughly tied | 92.91 points | Gemini 3 Flash (preview) Sent straight to the model | 92.62 points |
Where the leader changes
| Part of a test | Leader there | Score |
|---|---|---|
| DocuBench · Photo (JPEG) | Gemini 3.5 Flash Sent straight to the model | 95.2% |
| DocuBench · Hebrew | Gemini 3.5 Flash Sent straight to the model | 82.5% |
| ExtractBench (LlamaIndex) · Short documents | GPT-6 Astra Sent straight to the model | 97.2% |
| ExtractBench (LlamaIndex) · Long documents | Claude Opus 4.8 Coding agent (Claude Code)roughly tied, 3 setupsAlso roughly tied: GPT-5.5 Coding agent (Codex), asked to show where each value came from 88.0%; GPT-6 Sol Coding agent (Codex), asked to show where each value came from 87.3% | 88.1% |
| ParseBench · Tables | Claude Opus 5.5 Sent straight to the model, effort highroughly tiedAlso roughly tied: Claude Opus 5.5 Sent straight to the model, default setting 93.86 points | 94.25 points |
| ParseBench · Keeping the text complete and correct | Claude Opus 5.5 Sent straight to the model, default settingroughly tied, 5 setupsAlso roughly tied: Claude Opus 5.5 Sent straight to the model, effort high 91.72 points; Claude Opus 5.5 Sent straight to the model, effort low 91.68 points; Claude Fable 5.1 Sent straight to the model, default setting 91.19 points; Gemini 3 Flash (preview) Sent straight to the model, thinking high 90.87 points | 91.81 points |
| ParseBench · Formatting that carries meaning | Gemini 3.5 Flash Sent straight to the model, thinking highroughly tied, 4 setupsAlso roughly tied: GPT-5.6 Sol Sent straight to the model, reasoning high 77.16 points; Claude Opus 5.5 Sent straight to the model, default setting 77.04 points; Claude Opus 5.5 Sent straight to the model, effort high 77.04 points | 77.85 points |
| ParseBench · Placing each part on the page | Gemini 3.8 Flash Sent straight to the model, thinking highroughly tiedAlso roughly tied: Gemini 3.8 Flash Sent straight to the model, thinking low 72.24 points | 72.91 points |
| IDP Core Bench · Pulling out fields | Gemini 3 Flash (preview) Sent straight to the model, default settings | 91.1% |
| IDP Core Bench · Reading text | Gemini 3 Pro (preview) Sent straight to the model, default settingsroughly tiedAlso roughly tied: Gemini 3 Flash (preview) Sent straight to the model, default settings 81.7% | 81.8% |
| IDP Core Bench · Reading tables | Claude Sonnet 4.6 Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: Claude Opus 4.6 Sent straight to the model, default settings 96.0%; Gemini 3 Pro (preview) Sent straight to the model, default settings 95.8% | 96.3% |
| IDP Core Bench · Answering questions about a document | Claude Sonnet 4.6 Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: GPT-5 mini Sent straight to the model, default settings 65.0%; Claude Opus 4.6 Sent straight to the model, default settings 64.4% | 65.2% |
| Model | DocuBench | ExtractBench (LlamaIndex) | ParseBench | IDP Core Bench | OmniDocBench |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | 24 of 26Sent straight to the model, thinking on | 8 of 9Sent straight to the model, default settings | |||
| Claude Opus 4.6 | 23 of 26Sent straight to the model, default setting | 3 of 9Sent straight to the model, default settings | |||
| Claude Opus 4.8 | 8 of 12Coding agent (Claude Code) | 17 of 26Sent straight to the model, default setting | |||
| Claude Opus 5.5 | 9 of 12Coding agent (Claude Code), asked to show where each value came from | 1 of 26Sent straight to the model, effort high | |||
| Claude Sonnet 5 | 1 of 3Sent straight to the model | 20 of 26Sent straight to the model, default setting | |||
| Gemini 3 Flash (preview) | 4 of 26Sent straight to the model, thinking high | 4 of 9Sent straight to the model, default settings | 2 of 3Sent straight to the model | ||
| Gemini 3 Pro (preview) | 1 of 9Sent straight to the model, default settings | 1 of 3Sent straight to the model | |||
| Gemini 3.5 Flash | 3 of 3Sent straight to the model | 11 of 12Sent straight to the model | 5 of 26Sent straight to the model, thinking high | ||
| Gemini 3.8 Flash | 10 of 12Sent straight to the model | 6 of 26Sent straight to the model, thinking high | |||
| GPT-5 mini | 25 of 26Sent straight to the model, reasoning medium | 7 of 9Sent straight to the model, default settings | |||
| GPT-5.2 | 5 of 9Sent straight to the model, default settings | 3 of 3Sent straight to the model | |||
| GPT-5.4 Nano | 12 of 12Sent straight to the model | 26 of 26Sent straight to the model, default setting | |||
| GPT-5.5 | 2 of 3Sent straight to the model | 3 of 12Coding agent (Codex) | 15 of 26Sent straight to the model, reasoning medium | ||
| GPT-5.6 Luna | 6 of 12Coding agent (Codex), asked to show where each value came from | 14 of 26Sent straight to the model, reasoning max | |||
| GPT-5.6 Sol | 2 of 12Coding agent (Codex), asked to show where each value came from | 3 of 26Sent straight to the model, reasoning high | |||
| GPT-5.6 Terra | 4 of 12Coding agent (Codex), asked to show where each value came from | 13 of 26Sent straight to the model, reasoning medium | |||
| GPT-6 Astra | 5 of 12Sent straight to the model | 9 of 26Sent straight to the model, reasoning low | |||
| GPT-6 Luna | 7 of 12Coding agent (Codex), asked to show where each value came from | 16 of 26Sent straight to the model, reasoning max | |||
| GPT-6 Sol | 1 of 12Coding agent (Codex), asked to show where each value came from | 11 of 26Sent straight to the model, reasoning high |
The tests here put 34 pairs of models in a different order.
- In ParseBench, Claude Haiku 4.5, sent straight to the model, thinking on, placed above GPT-5 mini, sent straight to the model, reasoning medium; in IDP Core Bench, GPT-5 mini, sent straight to the model, default settings, placed above Claude Haiku 4.5, sent straight to the model, default settings.
- In ParseBench, Gemini 3 Flash (preview), sent straight to the model, thinking high, placed above Claude Opus 4.6, sent straight to the model, default setting; in IDP Core Bench, Claude Opus 4.6, sent straight to the model, default settings, placed above Gemini 3 Flash (preview), sent straight to the model, default settings.
- In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, default setting.
Show all 34 pairs placed in a different order
- In ParseBench, Claude Haiku 4.5, sent straight to the model, thinking on, placed above GPT-5 mini, sent straight to the model, reasoning medium; in IDP Core Bench, GPT-5 mini, sent straight to the model, default settings, placed above Claude Haiku 4.5, sent straight to the model, default settings.
- In ParseBench, Gemini 3 Flash (preview), sent straight to the model, thinking high, placed above Claude Opus 4.6, sent straight to the model, default setting; in IDP Core Bench, Claude Opus 4.6, sent straight to the model, default settings, placed above Gemini 3 Flash (preview), sent straight to the model, default settings.
- In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, default setting.
- In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Claude Opus 4.8, sent straight to the model, default setting.
- In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above Claude Opus 4.8, sent straight to the model, default setting.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, reasoning high.
- In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
- In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
- In DocuBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Claude Sonnet 5, sent straight to the model, default setting.
- In DocuBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.5, sent straight to the model; in ParseBench, GPT-5.5, sent straight to the model, reasoning medium, placed above Claude Sonnet 5, sent straight to the model, default setting.
- In ExtractBench (LlamaIndex), Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Gemini 3.8 Flash, sent straight to the model, thinking high.
- In DocuBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
- In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
- In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
- In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Luna, sent straight to the model, reasoning max, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Terra, sent straight to the model, reasoning medium, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-5.5, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Sol, sent straight to the model, reasoning high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
- In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
- In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-6 Sol, sent straight to the model, reasoning high.
Each test is listed on its own.
DocuBench
58 documents
Most accurate
- Claude Sonnet 5 Sent straight to the model92.3%
- GPT-5.5 Sent straight to the model72.5%
- Gemini 3.5 Flash Sent straight to the model67.0%
No cost published
ExtractBench (LlamaIndex)
370 documents
Most accurate
- GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
- GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
- GPT-5.5 Coding agent (Codex)93.6%
- GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
- GPT-6 Astra Sent straight to the model91.9%
- GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
- GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
- Claude Opus 4.8 Coding agent (Claude Code)87.1%
What this test found for PDF
- In ExtractBench (LlamaIndex), every model and setup scored lower on long documents than on short ones, by 2.0 to 65.5 points (14 compared). See the evidence
- The authors of ExtractBench (LlamaIndex) report that when a whole long document goes to a model in one request, the model often cuts long lists short. See the evidence
- On long documents in ExtractBench (LlamaIndex), the best agent setup scored 88.1% and the best single request scored 35.8%, a gap of 52.3 points. See the evidence
- On short documents in ExtractBench (LlamaIndex), asking the model to show where each value came from changed accuracy by at most 0.5 points (2 models tried both ways). See the evidence
- On long documents it was different: asking for sources changed accuracy by as much as 10.1 points, for Claude Opus 4.8, coding agent (Claude Code); and GPT-5.5, coding agent (Codex). See the evidence
OmniDocBench
PDF pages sent to the model as images
Most accurate
- Gemini 3 Pro (preview) Sent straight to the model92.91 points
- Gemini 3 Flash (preview) Sent straight to the model92.62 points
- GPT-5.2 Sent straight to the model86.59 points
No cost published
DocuBench
5 documents
Most accurate
- Gemini 3.5 Flash Sent straight to the model95.2%
- GPT-5.5 Sent straight to the model88.3%
- Claude Sonnet 5 Sent straight to the model75.9%
What this test found for JPEG
- On the image files (JPEG) in DocuBench, Gemini 3.5 Flash, sent straight to the model, led with 95.2%, while Claude Sonnet 5, sent straight to the model, led the whole test with 91.7%. The best model for one kind of document may not be the best for yours. See the evidence
No cost published
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
DocuBench
Tested on 2 documents, too few to rank Tester's page for DocuBench
DocuBench
Tested on 1 document, too few to rank Tester's page for DocuBench
Quality and cost
In ExtractBench (LlamaIndex) only. Cost as published by LlamaIndex. Your costs will differ.
Most accurate, and the cheapest within 1 point of the best: GPT-6 Sol Coding agent (Codex), asked to show where each value came from 94.6% · $0.098 per page
Best at each price
- 1GPT-5.4 Nano Sent straight to the model 74.9% · $0.0021 per page
- 2Gemini 3.8 Flash Sent straight to the model 80.7% · $0.0043 per page
- 3GPT-6 Luna Coding agent (Codex), asked to show where each value came from 91.0% · $0.0097 per page
- 4GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from 91.2% · $0.012 per page
- 5GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from 92.3% · $0.0945 per page
- 6GPT-6 Sol Coding agent (Codex), asked to show where each value came from 94.6% · $0.098 per pagebest quality
In ParseBench only. Cost as published by LlamaIndex. Your costs will differ.
Most accurate, and the cheapest within 1 point of the best: Claude Opus 5.5 Sent straight to the model, effort high 79.85 points · $0.0613 per page
Best at each price
- 1GPT-6 Luna Sent straight to the model, reasoning none 52.93 points · $0.0009 per page
- 2GPT-6 Luna Sent straight to the model, reasoning low 59.32 points · $0.0011 per page
- 3GPT-6 Luna Sent straight to the model, reasoning medium 62.38 points · $0.0014 per page
- 4GPT-6 Luna Sent straight to the model, reasoning high 64.21 points · $0.0021 per page
- 5GPT-6 Luna Sent straight to the model, reasoning xhigh 64.32 points · $0.0029 per page
- 6GPT-5.6 Luna Sent straight to the model, reasoning high 67.37 points · $0.0059 per page
- 7Gemini 3 Flash (preview) Sent straight to the model, thinking minimal 71.04 points · $0.0065 per page
- 8Gemini 3.7 Flash Sent straight to the model, thinking high 71.27 points · $0.0191 per page
- 9Gemini 3 Flash (preview) Sent straight to the model, thinking medium 75 points · $0.0198 per page
- 10Gemini 3 Flash (preview) Sent straight to the model, thinking high 75.05 points · $0.0241 per page
- 11Claude Opus 5.5 Sent straight to the model, default setting 78.01 points · $0.0579 per page
- 12Claude Opus 5.5 Sent straight to the model, effort high 79.85 points · $0.0613 per pagebest quality
Where it goes wrong
- F1measuredIn ExtractBench (LlamaIndex), every model and setup scored lower on long documents than on short ones, by 2.0 to 65.5 points (14 compared).evidence
- F2authors reportThe authors of ExtractBench (LlamaIndex) report that when a whole long document goes to a model in one request, the model often cuts long lists short.evidence
Technical details for this finding
“Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost.”
ExtractBench (LlamaIndex) - F3measuredOn the image files (JPEG) in DocuBench, Gemini 3.5 Flash, sent straight to the model, led with 95.2%, while Claude Sonnet 5, sent straight to the model, led the whole test with 91.7%. The best model for one kind of document may not be the best for yours.evidence
- F4measuredOn the Hebrew documents in DocuBench, Gemini 3.5 Flash, sent straight to the model, led with 82.5%, while Claude Sonnet 5, sent straight to the model, led the whole test with 91.7%.evidence
What helps
- F5measuredOn long documents in ExtractBench (LlamaIndex), the best agent setup scored 88.1% and the best single request scored 35.8%, a gap of 52.3 points.evidence
- F6measuredOn short documents in ExtractBench (LlamaIndex), asking the model to show where each value came from changed accuracy by at most 0.5 points (2 models tried both ways).evidence
- F7measuredOn long documents it was different: asking for sources changed accuracy by as much as 10.1 points, for Claude Opus 4.8, coding agent (Claude Code); and GPT-5.5, coding agent (Codex).evidence
- P1common practiceBefore you rely on a model, test it on a sample of your own documents and compare its answers with ones you have already checked.Sources: UK Government, AI Playbook for the UK Government; NIST, AI RMF Playbook, Measure; Anthropic, Define success criteria and build evaluations
- P2common practiceHave a person review any figure you will pay, file or sign before it goes out.Sources: UK Government, AI Playbook for the UK Government; Microsoft, Interpret and improve accuracy and confidence scores (Document Intelligence)
- P3common practiceTest image files, such as photos and scans, and documents in other languages separately from clean files.Sources: NIST, AI RMF Playbook, Measure; OpenAI, Evaluation best practices
Next steps for a small business
- Test on a sample of your own documents and compare the answers with ones you have already checked.P1 behind step 1: Test on a sample of your own documents and compare the answers with ones you have already checked.
- For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.F5 behind step 2: For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.F2 behind step 2: For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.
- If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.F6 behind step 3: If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.F7 behind step 3: If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.
- Keep a person reviewing anything you will pay, file or sign.P2 behind step 4: Keep a person reviewing anything you will pay, file or sign.
- If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.F3 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.F4 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.P3 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.
F: a finding above. P: a common practice above.
Everyday jobs, each matched to the test that is most like it. This is simple arithmetic on the test's score. Your documents will differ.
Invoices for accounts payable
Pulling 12 fields, such as vendor, invoice number, date and total, from each of 500 invoices a month: 6,000 fields. · DocuBenchbest: about 1 in 12 wronglowest: about 1 in 4
Pulling 12 fields, such as vendor, invoice number, date and total, from each of 500 invoices a month: 6,000 fields.
Closest test: DocuBench. DocuBench scores each field as right or wrong, and its documents include invoices, statements and bills like the ones an office handles.
- Claude Sonnet 5, sent straight to the model (91.7%): about 1 in 12 fields wrong, roughly 500 of 6,000 a month.
- Gemini 3.5 Flash, sent straight to the model (73.0%): about 1 in 4 fields wrong, roughly 1,600 of 6,000 a month.
Phone photos of receipts
Pulling 10 fields, such as store, date and total, from each of 100 receipts photographed with a phone: 1,000 fields a month. · DocuBench, Photo (JPEG)best: about 1 in 21 wronglowest: about 1 in 4
Pulling 10 fields, such as store, date and total, from each of 100 receipts photographed with a phone: 1,000 fields a month.
Closest test: DocuBench, Photo (JPEG) only. This DocuBench group covers five photo (JPEG) files, including one receipt photo. It is a small, mixed group, so it is only a loose match for phone photos of receipts.
- Gemini 3.5 Flash, sent straight to the model (95.2%): about 1 in 21 fields wrong, roughly 48 of 1,000 a month.
- Claude Sonnet 5, sent straight to the model (75.9%): about 1 in 4 fields wrong, roughly 240 of 1,000 a month.
Long statements and registers
Pulling every line from long documents, such as a year of account statements or a check register that runs past 50 pages, where missing even a few lines matters. · ExtractBench (LlamaIndex), Long documentsOpen for the figures
Pulling every line from long documents, such as a year of account statements or a check register that runs past 50 pages, where missing even a few lines matters.
Closest test: ExtractBench (LlamaIndex), Long documents only. The long documents in ExtractBench (LlamaIndex) each run more than 50 pages, like its example check register with 3,308 payment lines.
- Claude Opus 4.8, coding agent (Claude Code): 88.1%
- GPT-5.5, coding agent (Codex), asked to show where each value came from: 88.0%
- GPT-6 Sol, coding agent (Codex), asked to show where each value came from: 87.3%
- GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from: 82.7%
- GPT-6 Luna, coding agent (Codex), asked to show where each value came from: 80.9%
- Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from: 80.8%
- GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from: 80.0%
- GPT-5.5, coding agent (Codex): 78.9%
- Claude Opus 4.8, coding agent (Claude Code), asked to show where each value came from: 77.9%
- GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from: 77.4%
- GPT-5.4 Nano, sent straight to the model: 35.8%
- GPT-6 Astra, sent straight to the model: 31.7%
- Gemini 3.8 Flash, sent straight to the model: 28.7%
- Gemini 3.5 Flash, sent straight to the model: 27.9%
This score is not a simple share of correct answers, so we do not turn it into a count.
DocuBench
best 91.7%DocuPipe72 documentsNewest result 2026-07-17reliability not measuredcaveat
A test of how well AI pulls requested information out of 72 hard documents, checked against answers that people verified by hand.
Results
All Sent straight to the model · Run date not published; posted 2026-07-17
- Claude Sonnet 591.7%
- GPT-5.576.5%
- Gemini 3.5 Flash73.0%
Watch out
Most of the test is in English: 47 of the 72 documents. Several languages have only one to three documents each, so treat their results as examples. The authors say 72 documents are enough for careful inspection but not for broad claims about every kind of document.
Reliability not measured by this source.
DocuPipe built this benchmark and sells a document extraction product, which it tests alongside the general models. Only the general models are shown here.
See DocuPipe's resultsNot tested here: 28 models
Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.2, GPT-5.4, GPT-5.4 Nano, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: DocuBench
What's in the test
Each test item is one document, a list of the fields to pull out of it, and the correct answer for each field. The documents come in 10 file types, from PDFs and photos to spreadsheets and web pages, and in 12 languages. Most come from public sources such as government publications and sample documents that companies published; a few were created by the benchmark's authors.
72 documents
- Invoices
- Statements
- Utility bills
- Annual reports
- Payslips
- Purchase orders
- Shipping waybills
- Healthcare forms
- Engineering drawings
- Insurance declarations
- Dictionaries
- Directories
- Auction catalogs
- Government registers
- Spreadsheets
An example
CPS Energy electric bill
A two-page home electric bill that the utility published as a sample, with personal details blacked out. The model must pull out the billing period, the meter readings, each charge line and the total, and must not count the total row as a charge.
- Billing period:
- September 18, 2021 to October 18, 2021
- Current meter reading:
- 48,374
- Electricity used:
- 463 kilowatt-hours
- One charge line:
- Service Availability Charge, $8.75
- Total electric bill:
- $58.35
Where results change
By kind of file
The same test split by the kind of file the document came in.
| Model and setup | Photo (JPEG) | |
|---|---|---|
| Claude Sonnet 5Sent straight to the model | 75.9% | 92.3% |
| GPT-5.5Sent straight to the model | 88.3% | 72.5% |
| Gemini 3.5 FlashSent straight to the model | 95.2% | 67.0% |
By language
The same test split by the document's language.
| Model and setup | English | Hebrew |
|---|---|---|
| Claude Sonnet 5Sent straight to the model | 95.2% | 70.6% |
| GPT-5.5Sent straight to the model | 80.2% | 76.5% |
| Gemini 3.5 FlashSent straight to the model | 75.1% | 82.5% |
Technical details for DocuBench
Macro-average field accuracy: The average share of fields pulled out correctly per document. A failed request counts every field on that document as missed.
- Claude Sonnet 5, sent straight to the model: Direct API call, one request per document. Tools: None. Input: PDF sent as a document. Full output budget, per the source's README.
- GPT-5.5, sent straight to the model: Direct API call, one request per document. Tools: None. Input: File upload. Full output budget, per the source's README.
- Gemini 3.5 Flash, sent straight to the model: Direct API call, one request per document. Tools: None. Input: File sent inline. Full output budget, per the source's README.
Source version: DocuBench 0.2.0
Results posted 2026-07-17
Run dates not published by the tester; the dates are when results were posted.
License: CC BY 4.0 (results folder). Checked 2026-09-25.
Based on one or a few documents; treat as an example, not a pattern.
- By kind of file, CSV file (1 document): Claude Sonnet 5, sent straight to the model: 97.1%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, Word document (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 83.3%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, Web page (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, Image (PNG) (1 document): Claude Sonnet 5, sent straight to the model: 95.7%; GPT-5.5, sent straight to the model: 79.4%; Gemini 3.5 Flash, sent straight to the model: 92.2%
- By kind of file, Scan (TIFF) (1 document): Claude Sonnet 5, sent straight to the model: 85.7%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, Text file (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, Spreadsheet (Excel) (1 document): Claude Sonnet 5, sent straight to the model: 93.2%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By kind of file, XML file (2 documents): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By language, Arabic (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By language, German (2 documents): Claude Sonnet 5, sent straight to the model: 96.1%; GPT-5.5, sent straight to the model: 96.2%; Gemini 3.5 Flash, sent straight to the model: 99.6%
- By language, Spanish (3 documents): Claude Sonnet 5, sent straight to the model: 98.3%; GPT-5.5, sent straight to the model: 60.7%; Gemini 3.5 Flash, sent straight to the model: 35.8%
- By language, French (1 document): Claude Sonnet 5, sent straight to the model: 81.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 94.8%
- By language, Hindi (2 documents): Claude Sonnet 5, sent straight to the model: 85.8%; GPT-5.5, sent straight to the model: 8.5%; Gemini 3.5 Flash, sent straight to the model: 8.2%
- By language, Italian (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 96.5%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By language, Japanese (4 documents): Claude Sonnet 5, sent straight to the model: 79.7%; GPT-5.5, sent straight to the model: 48.1%; Gemini 3.5 Flash, sent straight to the model: 52.9%
- By language, Dutch (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 76.7%; Gemini 3.5 Flash, sent straight to the model: 100.0%
- By language, Portuguese (1 document): Claude Sonnet 5, sent straight to the model: 96.4%; GPT-5.5, sent straight to the model: 97.6%; Gemini 3.5 Flash, sent straight to the model: 96.4%
- By language, Chinese (2 documents): Claude Sonnet 5, sent straight to the model: 90.2%; GPT-5.5, sent straight to the model: 72.5%; Gemini 3.5 Flash, sent straight to the model: 59.7%
Sources for this summary:
ExtractBench (LlamaIndex)
best 94.6%LlamaIndex370 documentsNewest result 2026-09-23reliability not measuredcaveat
A test of how well AI fills in a set list of fields from 370 business documents, from short forms to reports of more than 50 pages, checked field by field against the correct answers.
Results
Run date not published; posted 2026-09-23
Each model's best setting in this test
- GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
- GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
- GPT-5.5 Coding agent (Codex)93.6%
- GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
- GPT-6 Astra Sent straight to the model91.9%
- GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
- GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
- Claude Opus 4.8 Coding agent (Claude Code)87.1%
Show all 14 results (6 not shown above) from ExtractBench (LlamaIndex)
- GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
- GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
- GPT-5.5 Coding agent (Codex)93.6%
- GPT-5.5 Coding agent (Codex), asked to show where each value came from93.3%
- GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
- GPT-6 Astra Sent straight to the model91.9%
- GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
- GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
- Claude Opus 4.8 Coding agent (Claude Code)87.1%
- Claude Opus 4.8 Coding agent (Claude Code), asked to show where each value came from86.8%
- Claude Opus 5.5 Coding agent (Claude Code), asked to show where each value came from86.1%
- Gemini 3.8 Flash Sent straight to the model80.7%
- Gemini 3.5 Flash Sent straight to the model79.8%
- GPT-5.4 Nano Sent straight to the model74.9%
Watch out
The leaderboard does not say when each model was run; the date shown is when the leaderboard was last updated. The authors report that models sent a whole document in one direct call do well on short documents but often cut long lists short, while coding agents stay more accurate at much higher cost.
Reliability not measured by this source.
LlamaIndex built this benchmark and sells a document extraction product, LlamaExtract, which it tests alongside the general models. Only the general models are shown here.
See LlamaIndex's resultsNot tested here: 19 models
Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.2, GPT-5.4.
More about this test: ExtractBench (LlamaIndex)
What's in the test
The documents come from public records such as company and regulatory filings, government purchasing and customs forms, court exhibits and Texas energy filings. Of the 370, 45 are long lists created by the benchmark's authors from real layouts. The test sorts them by length: 252 short (10 pages or fewer), 98 medium (11 to 50 pages) and 20 long (more than 50 pages).
370 documents
- Finance and fund holdings
- Energy regulatory forms
- Government purchasing and customs forms
- Auto valuation
- Supply chain
- Healthcare payment notices
- Legal and bankruptcy filings
- Real estate
An example
Freer ISD check register
A real report from the Freer Independent School District in Texas listing every check the district wrote in one school year. It runs 123 pages. The model must list every payment line and copy each check's number, date and payee onto the lines below it, where they are not printed again.
- District name:
- Freer ISD
- Report title:
- YTD Check Register
- Report period:
- September 1, 2022 to August 31, 2023
- Total amount:
- $5,991,776.11
- Payment lines to find:
- 3,308
Where results change
By document length
The same test split by how long the documents are.
| Model and setup | Short documents | Medium documents | Long documents |
|---|---|---|---|
| GPT-6 SolCoding agent (Codex), asked to show where each value came from | 96.0% | 92.4% | 87.3% |
| GPT-5.6 SolCoding agent (Codex), asked to show where each value came from | 96.0% | 90.2% | 82.7% |
| GPT-5.5Coding agent (Codex) | 95.7% | 91.2% | 78.9% |
| GPT-5.5Coding agent (Codex), asked to show where each value came from | 95.6% | 88.7% | 88.0% |
| GPT-5.6 TerraCoding agent (Codex), asked to show where each value came from | 95.6% | 86.7% | 77.4% |
| GPT-6 AstraSent straight to the model | 97.2% | 90.6% | 31.7% |
| GPT-5.6 LunaCoding agent (Codex), asked to show where each value came from | 93.5% | 87.4% | 80.0% |
| GPT-6 LunaCoding agent (Codex), asked to show where each value came from | 93.1% | 87.7% | 80.9% |
| Claude Opus 4.8Coding agent (Claude Code) | 90.1% | 79.2% | 88.1% |
| Claude Opus 4.8Coding agent (Claude Code), asked to show where each value came from | 90.5% | 79.2% | 77.9% |
| Claude Opus 5.5Coding agent (Claude Code), asked to show where each value came from | 89.6% | 78.0% | 80.8% |
| Gemini 3.8 FlashSent straight to the model | 88.1% | 72.2% | 28.7% |
| Gemini 3.5 FlashSent straight to the model | 87.9% | 69.8% | 27.9% |
| GPT-5.4 NanoSent straight to the model | 77.4% | 76.4% | 35.8% |
Technical details for ExtractBench (LlamaIndex)
Unified value F1: An F1 score for extracted values, counting every item in a list. It rewards values that match the correct answer, allowing for small formatting differences such as how dates are written. It penalizes both wrong and missing values, and it is averaged evenly across documents. Rows marked Evidence are runs where the agent was also asked to cite the page and location behind each value it extracted.
- GPT-6 Sol, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- GPT-5.5, coding agent (Codex): Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: standard.
- GPT-5.5, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
- GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- GPT-6 Luna, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- Claude Opus 4.8, coding agent (Claude Code): Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: standard.
- Claude Opus 4.8, coding agent (Claude Code), asked to show where each value came from: Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from: Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
- Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
- Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
- GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
Source version: ExtractBench leaderboard, 2026-09-23
Results posted 2026-09-23
Run dates not published by the tester; the dates are when results were posted.
License: Apache-2.0; LlamaIndex confirmed public use of the results by email on 2026-09-24. Checked 2026-09-26.
ParseBench
best 79.85 pointsLlamaIndex2,078 pagesNewest result 2026-09-25reliability not measuredcaveat
A test from LlamaIndex of how well AI turns document pages into clean, structured text: tables, charts, the words on the page, the formatting that carries meaning, and where each part sits on the page.
Results
Each model's best setting in this test
- Claude Opus 5.5 Sent straight to the model, effort high, Run date not published; posted 2026-09-2479.85 points
- Claude Fable 5.1 Sent straight to the model, default setting, Run date not published; posted 2026-09-0278.92 points
- GPT-5.6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2575.35 points
- Gemini 3 Flash (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-04-2175.05 points
- Gemini 3.5 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2574.02 points
- Gemini 3.8 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-0472.08 points
- Gemini 3.7 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2571.27 points
- Claude Fable 5 Sent straight to the model, default setting, Run date not published; posted 2026-06-1070.78 points
Show all 61 results (53 not shown above) from ParseBench
- Claude Opus 5.5 Sent straight to the model, effort high, Run date not published; posted 2026-09-2479.85 points
- Claude Fable 5.1 Sent straight to the model, default setting, Run date not published; posted 2026-09-0278.92 points
- Claude Opus 5.5 Sent straight to the model, default setting, Run date not published; posted 2026-09-2278.01 points
- GPT-5.6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2575.35 points
- Gemini 3 Flash (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-04-2175.05 points
- Gemini 3 Flash (preview) Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2575 points
- Claude Opus 5.5 Sent straight to the model, effort low, Run date not published; posted 2026-09-2474.47 points
- Gemini 3.5 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2574.02 points
- Gemini 3.8 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-0472.08 points
- Gemini 3.7 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2571.27 points
- Gemini 3 Flash (preview) Sent straight to the model, thinking minimal, Run date not published; posted 2026-04-2171.04 points
- Claude Fable 5 Sent straight to the model, default setting, Run date not published; posted 2026-06-1070.78 points
- Gemini 3.7 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2570.78 points
- Gemini 3.8 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2570.7 points
- GPT-6 Astra Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2470.67 points
- Gemini 3.8 Flash Sent straight to the model, thinking low, Run date not published; posted 2026-09-0470.15 points
- Gemini 3.6 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2570.03 points
- Gemini 3.5 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-06-0169.92 points
- GPT-6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2569.69 points
- Gemini 3.7 Flash Sent straight to the model, thinking low, Run date not published; posted 2026-09-2569.57 points
- Gemini 3.1 Pro (preview) Sent straight to the model, default setting, Run date not published; posted 2026-04-2169.14 points
- GPT-5.6 Terra Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2568.39 points
- GPT-5.6 Luna Sent straight to the model, reasoning max, Run date not published; posted 2026-09-2468.34 points
- GPT-5.6 Luna Sent straight to the model, reasoning xhigh, Run date not published; posted 2026-09-2468.24 points
- GPT-6 Sol Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2268.19 points
- GPT-5.6 Terra Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2568.07 points
- GPT-5.5 Sent straight to the model, reasoning medium, Run date not published; posted 2026-04-2467.76 points
- GPT-5.6 Luna Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2467.37 points
- Gemini 3.6 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-08-0466.78 points
- GPT-6 Sol Sent straight to the model, reasoning none, Run date not published; posted 2026-09-2266.1 points
- GPT-6 Luna Sent straight to the model, reasoning max, Run date not published; posted 2026-09-2465.77 points
- Gemini 3.6 Flash Sent straight to the model, thinking minimal, Run date not published; posted 2026-08-0465.6 points
- GPT-5.5 Sent straight to the model, reasoning none, Run date not published; posted 2026-04-2464.39 points
- GPT-6 Luna Sent straight to the model, reasoning xhigh, Run date not published; posted 2026-09-2464.32 points
- GPT-6 Luna Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2464.21 points
- GPT-5.6 Terra Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1364.17 points
- Claude Opus 4.8 Sent straight to the model, default setting, Run date not published; posted 2026-06-0163.7 points
- GPT-5.6 Luna Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2463.58 points
- Claude Opus 4.7 Sent straight to the model, default setting, Run date not published; posted 2026-04-2163.34 points
- Gemini 3.5 Flash Sent straight to the model, thinking minimal, Run date not published; posted 2026-06-0163.09 points
- GPT-6 Luna Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2262.38 points
- GPT-5.4 Sent straight to the model, reasoning none, Run date not published; posted 2026-04-2162.23 points
- Claude Sonnet 5 Sent straight to the model, default setting, Run date not published; posted 2026-07-0262.13 points
- GPT-5.6 Sol Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1362.12 points
- Gemini 3.1 Pro (preview) Sent straight to the model, thinking low, Run date not published; posted 2026-09-2560.82 points
- Gemini 3.5 Flash-Lite Sent straight to the model, thinking high, Run date not published; posted 2026-09-2560.64 points
- GPT-5.6 Luna Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2460.18 points
- GPT-6 Luna Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2459.32 points
- Gemini 3.1 Flash-Lite (preview) Sent straight to the model, default setting, Run date not published; posted 2026-04-2158.32 points
- Gemini 3.1 Flash-Lite (preview) Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2558.13 points
- Gemini 3.1 Flash-Lite (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-09-2557.29 points
- Gemini 3.5 Flash-Lite Sent straight to the model, default setting, Run date not published; posted 2026-08-0457.15 points
- GPT-5.6 Luna Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1356.32 points
- Gemini 3.5 Flash-Lite Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2554.3 points
- Claude Opus 4.6 Sent straight to the model, default setting, Run date not published; posted 2026-04-2154.07 points
- Claude Haiku 4.5 Sent straight to the model, thinking on, Run date not published; posted 2026-04-2153.12 points
- GPT-6 Luna Sent straight to the model, reasoning none, Run date not published; posted 2026-09-2252.93 points
- GPT-5 mini Sent straight to the model, reasoning medium, Run date not published; posted 2026-04-2151.52 points
- GPT-5 mini Sent straight to the model, reasoning minimal, Run date not published; posted 2026-04-2146.83 points
- Claude Haiku 4.5 Sent straight to the model, thinking off, Run date not published; posted 2026-04-2145.17 points
- GPT-5.4 Nano Sent straight to the model, default setting, Run date not published; posted 2026-04-2143.35 points
Watch out
LlamaIndex does not publish when it ran each model. The date shown is the last time LlamaIndex changed that model's row in its results file, which can be later than the run, for example when it reprices a row. Each setting LlamaIndex tried is its own row, so one model can appear up to six times. LlamaIndex works out each cost from the tokens that run used, priced at the model maker's list price when the row was last priced; when prices change it reprices the row without running it again, so a cost can change while the score stays the same. Each model is run once, one page at a time; LlamaIndex publishes no repeat runs. An Anthropic employee contributed three changes to ParseBench's code on 2026-08-04: two to how formatting is scored and one option that none of the leaderboard's runs use. Every change to every row shown was made by LlamaIndex.
Reliability not measured by this source.
LlamaIndex built this benchmark and sells a document parsing product, LlamaParse, which it tests alongside the general models. LlamaParse held the top place on the leaderboard when we checked on 2026-09-29. Only the general models are shown here.
See LlamaIndex's resultsNot tested here: 5 models
Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-4.1, GPT-5 nano, GPT-5.2.
More about this test: ParseBench
What's in the test
About 2,000 pages checked by people, from more than 1,200 public documents in insurance, finance, government and other fields. Most pages are PDF; 42, all in the page-placement part, are JPG or PNG images, which are sent to the model as images. ParseBench publishes no separate scores for them.
2,078 pages
- Tables
- Charts
- Text pages
- Page layout
Where results change
By what is checked
The five scores the overall averages, each from 0 to 100 and each scored its own way.
| Model and setup | Tables | Charts | Keeping the text complete and correct | Formatting that carries meaning | Placing each part on the page |
|---|---|---|---|---|---|
| Claude Opus 5.5Sent straight to the model, effort high | 94.25 points | 70.89 points | 91.72 points | 77.04 points | 65.33 points |
| Claude Fable 5.1Sent straight to the model, default setting | 91.52 points | 67.06 points | 91.19 points | 76.52 points | 68.3 points |
| Claude Opus 5.5Sent straight to the model, default setting | 93.86 points | 64.12 points | 91.81 points | 77.04 points | 63.2 points |
| GPT-5.6 SolSent straight to the model, reasoning high | 91.39 points | 68.68 points | 88.16 points | 77.16 points | 51.35 points |
| Gemini 3 Flash (preview)Sent straight to the model, thinking high | 91.5 points | 64.79 points | 90.87 points | 68.31 points | 59.77 points |
| Gemini 3 Flash (preview)Sent straight to the model, thinking medium | 91.01 points | 61.56 points | 88.67 points | 67.95 points | 65.79 points |
| Claude Opus 5.5Sent straight to the model, effort low | 92.69 points | 53.52 points | 91.68 points | 75.21 points | 59.27 points |
| Gemini 3.5 FlashSent straight to the model, thinking high | 90.44 points | 45.27 points | 88.92 points | 77.85 points | 67.6 points |
| Gemini 3.8 FlashSent straight to the model, thinking high | 89.12 points | 32.34 points | 89.68 points | 76.36 points | 72.91 points |
| Gemini 3.7 FlashSent straight to the model, thinking high | 90.16 points | 32.07 points | 88.6 points | 74.98 points | 70.56 points |
| Gemini 3 Flash (preview)Sent straight to the model, thinking minimal | 89.85 points | 64.83 points | 86.19 points | 58.35 points | 55.97 points |
| Claude Fable 5Sent straight to the model, default setting | 89.79 points | 52.21 points | 90.02 points | 72.62 points | 49.24 points |
| Gemini 3.7 FlashSent straight to the model, thinking medium | 88.66 points | 35.11 points | 88.45 points | 73.05 points | 68.63 points |
| Gemini 3.8 FlashSent straight to the model, thinking medium | 89.08 points | 32.72 points | 88.42 points | 71.8 points | 71.5 points |
| GPT-6 AstraSent straight to the model, reasoning low | 93.17 points | 35.08 points | 89.6 points | 75.52 points | 59.99 points |
| Gemini 3.8 FlashSent straight to the model, thinking low | 88.18 points | 35.09 points | 88.25 points | 66.99 points | 72.24 points |
| Gemini 3.6 FlashSent straight to the model, thinking high | 89.3 points | 34.53 points | 88.21 points | 73.11 points | 64.98 points |
| Gemini 3.5 FlashSent straight to the model, thinking medium | 91.12 points | 44.01 points | 90.19 points | 59.14 points | 65.14 points |
| GPT-6 SolSent straight to the model, reasoning high | 92.63 points | 48.05 points | 87.81 points | 67.91 points | 52.06 points |
| Gemini 3.7 FlashSent straight to the model, thinking low | 87.96 points | 31.84 points | 88.44 points | 68.89 points | 70.71 points |
| Gemini 3.1 Pro (preview)Sent straight to the model, default setting | 91 points | 41.13 points | 90.16 points | 52.43 points | 70.99 points |
| GPT-5.6 TerraSent straight to the model, reasoning medium | 86.72 points | 67.86 points | 87.06 points | 68.47 points | 31.82 points |
| GPT-5.6 LunaSent straight to the model, reasoning max | 88.27 points | 49.16 points | 86.57 points | 64.95 points | 52.73 points |
| GPT-5.6 LunaSent straight to the model, reasoning xhigh | 89.42 points | 49.44 points | 87.13 points | 66.51 points | 48.69 points |
| GPT-6 SolSent straight to the model, reasoning medium | 90.27 points | 46.46 points | 87.3 points | 67.88 points | 49.05 points |
| GPT-5.6 TerraSent straight to the model, reasoning low | 86.91 points | 67.71 points | 86.83 points | 67.92 points | 30.97 points |
| GPT-5.5Sent straight to the model, reasoning medium | 90.05 points | 65.53 points | 86.81 points | 60.12 points | 36.28 points |
| GPT-5.6 LunaSent straight to the model, reasoning high | 88.75 points | 49.25 points | 86.99 points | 67.14 points | 44.74 points |
| Gemini 3.6 FlashSent straight to the model, thinking medium | 89.47 points | 31 points | 87.7 points | 58.46 points | 67.27 points |
| GPT-6 SolSent straight to the model, reasoning none | 89.49 points | 42.02 points | 86.38 points | 67.32 points | 45.27 points |
| GPT-6 LunaSent straight to the model, reasoning max | 87.42 points | 45.71 points | 84.44 points | 61.81 points | 49.48 points |
| Gemini 3.6 FlashSent straight to the model, thinking minimal | 85.5 points | 24.76 points | 83.3 points | 64.49 points | 69.94 points |
| GPT-5.5Sent straight to the model, reasoning none | 89.31 points | 59.11 points | 87.17 points | 64.46 points | 21.9 points |
| GPT-6 LunaSent straight to the model, reasoning xhigh | 88.01 points | 42.38 points | 85 points | 65.34 points | 40.86 points |
| GPT-6 LunaSent straight to the model, reasoning high | 87.97 points | 44.28 points | 85.27 points | 63.58 points | 39.97 points |
| GPT-5.6 TerraSent straight to the model, reasoning none | 86.23 points | 63.75 points | 82.56 points | 59.96 points | 28.35 points |
| Claude Opus 4.8Sent straight to the model, default setting | 89.65 points | 49.75 points | 89.02 points | 71.38 points | 18.69 points |
| GPT-5.6 LunaSent straight to the model, reasoning medium | 85.32 points | 46.91 points | 85.27 points | 65.8 points | 34.6 points |
| Claude Opus 4.7Sent straight to the model, default setting | 87.17 points | 55.84 points | 90.26 points | 69.42 points | 13.99 points |
| Gemini 3.5 FlashSent straight to the model, thinking minimal | 86.24 points | 18.74 points | 88.22 points | 68.09 points | 54.18 points |
| GPT-6 LunaSent straight to the model, reasoning medium | 86.74 points | 42.97 points | 84.94 points | 61.13 points | 36.14 points |
| GPT-5.4Sent straight to the model, reasoning none | 83.89 points | 65.22 points | 85.57 points | 59.52 points | 16.95 points |
| Claude Sonnet 5Sent straight to the model, default setting | 86.65 points | 60.53 points | 86.51 points | 64.93 points | 12.03 points |
| GPT-5.6 SolSent straight to the model, reasoning none | 89.34 points | 59.57 points | 86.17 points | 50.42 points | 25.09 points |
| Gemini 3.1 Pro (preview)Sent straight to the model, thinking low | 90.29 points | 11.43 points | 87.11 points | 45.56 points | 69.7 points |
| Gemini 3.5 Flash-LiteSent straight to the model, thinking high | 83.31 points | 19.27 points | 86.85 points | 54.03 points | 59.75 points |
| GPT-5.6 LunaSent straight to the model, reasoning low | 81.67 points | 38.95 points | 85.28 points | 64.75 points | 30.26 points |
| GPT-6 LunaSent straight to the model, reasoning low | 84.2 points | 39.05 points | 84.95 points | 59.9 points | 28.51 points |
| Gemini 3.1 Flash-Lite (preview)Sent straight to the model, default setting | 85.48 points | 9.92 points | 89.46 points | 58.38 points | 48.38 points |
| Gemini 3.1 Flash-Lite (preview)Sent straight to the model, thinking medium | 83.43 points | 14.35 points | 87.67 points | 56.17 points | 49.05 points |
| Gemini 3.1 Flash-Lite (preview)Sent straight to the model, thinking high | 76.76 points | 30.89 points | 83.28 points | 56.61 points | 38.93 points |
| Gemini 3.5 Flash-LiteSent straight to the model, default setting | 73.48 points | 6.26 points | 82.22 points | 57.74 points | 66.03 points |
| GPT-5.6 LunaSent straight to the model, reasoning none | 81.3 points | 28.32 points | 82.65 points | 63.9 points | 25.43 points |
| Gemini 3.5 Flash-LiteSent straight to the model, thinking medium | 72.85 points | 6.23 points | 78.8 points | 48.05 points | 65.55 points |
| Claude Opus 4.6Sent straight to the model, default setting | 86.52 points | 13.49 points | 89.7 points | 64.19 points | 16.47 points |
| Claude Haiku 4.5Sent straight to the model, thinking on | 78.7 points | 27.39 points | 84.83 points | 62.36 points | 12.3 points |
| GPT-6 LunaSent straight to the model, reasoning none | 81.42 points | 15.48 points | 85.07 points | 59.84 points | 22.85 points |
| GPT-5 miniSent straight to the model, reasoning medium | 74.6 points | 38.96 points | 85.68 points | 45.35 points | 13.03 points |
| GPT-5 miniSent straight to the model, reasoning minimal | 69.82 points | 30.13 points | 82.3 points | 45.77 points | 6.15 points |
| Claude Haiku 4.5Sent straight to the model, thinking off | 77.21 points | 13.77 points | 78.74 points | 49.39 points | 6.72 points |
| GPT-5.4 NanoSent straight to the model, default setting | 60.16 points | 17.57 points | 78.05 points | 54.56 points | 6.41 points |
Technical details for ParseBench
Overall: The average of five scores, each from 0 to 100: rebuilding tables, reading the numbers in charts, keeping the text complete and correct, keeping formatting that carries meaning such as headings, and placing each part of the page in the right spot.
- Claude Opus 5.5, sent straight to the model, effort high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: effort high.
- Claude Fable 5.1, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- Claude Opus 5.5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- GPT-5.6 Sol, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
- Gemini 3 Flash (preview), sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3 Flash (preview), sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- Claude Opus 5.5, sent straight to the model, effort low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: effort low.
- Gemini 3.5 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3.8 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3.7 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3 Flash (preview), sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
- Claude Fable 5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- Gemini 3.7 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- Gemini 3.8 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- GPT-6 Astra, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
- Gemini 3.8 Flash, sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
- Gemini 3.6 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3.5 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- GPT-6 Sol, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
- Gemini 3.7 Flash, sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
- Gemini 3.1 Pro (preview), sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- GPT-5.6 Terra, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- GPT-5.6 Luna, sent straight to the model, reasoning max: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning max.
- GPT-5.6 Luna, sent straight to the model, reasoning xhigh: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning xhigh.
- GPT-6 Sol, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- GPT-5.6 Terra, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
- GPT-5.5, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- GPT-5.6 Luna, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
- Gemini 3.6 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- GPT-6 Sol, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- GPT-6 Luna, sent straight to the model, reasoning max: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning max.
- Gemini 3.6 Flash, sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
- GPT-5.5, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- GPT-6 Luna, sent straight to the model, reasoning xhigh: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning xhigh.
- GPT-6 Luna, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
- GPT-5.6 Terra, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- Claude Opus 4.8, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- GPT-5.6 Luna, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- Claude Opus 4.7, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- Gemini 3.5 Flash, sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
- GPT-6 Luna, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- GPT-5.4, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- Claude Sonnet 5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- GPT-5.6 Sol, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- Gemini 3.1 Pro (preview), sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
- Gemini 3.5 Flash-Lite, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- GPT-5.6 Luna, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
- GPT-6 Luna, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
- Gemini 3.1 Flash-Lite (preview), sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- Gemini 3.1 Flash-Lite (preview), sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- Gemini 3.1 Flash-Lite (preview), sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
- Gemini 3.5 Flash-Lite, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- GPT-5.6 Luna, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- Gemini 3.5 Flash-Lite, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
- Claude Opus 4.6, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
- Claude Haiku 4.5, sent straight to the model, thinking on: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking on.
- GPT-6 Luna, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
- GPT-5 mini, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
- GPT-5 mini, sent straight to the model, reasoning minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning minimal.
- Claude Haiku 4.5, sent straight to the model, thinking off: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking off.
- GPT-5.4 Nano, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
Source version: ParseBench leaderboard.csv at commit afb36bd, downloaded 2026-09-29
Results posted 2026-04-21 to 2026-09-25
Run dates not published by the tester; the dates are when results were posted.
License: Apache-2.0; LlamaIndex offered it for public use by email on 2026-09-27. Checked 2026-09-29.
Sources for this summary:
IDP Core Bench
best 81.8%NanonetsNewest result 2026-03-23reliability not measuredcaveat
A test from Nanonets of how well AI reads everyday business documents: pulling out fields such as invoice numbers, dates and totals, reading printed and handwritten text, reading tables, and answering questions about a document.
Results
All Sent straight to the model, default settings · Run date not published; posted 2026-03-23
Each model's best setting in this test
- Gemini 3 Pro (preview)81.8%
- Claude Sonnet 4.681.2%
- Claude Opus 4.681.1%
- Gemini 3 Flash (preview)80.5%
- GPT-5.277.4%
- GPT-4.174.7%
- GPT-5 mini73.3%
- Claude Haiku 4.572.9%
Show all 9 results (1 not shown above) from IDP Core Bench
- Gemini 3 Pro (preview)81.8%
- Claude Sonnet 4.681.2%
- Claude Opus 4.681.1%
- Gemini 3 Flash (preview)80.5%
- GPT-5.277.4%
- GPT-4.174.7%
- GPT-5 mini73.3%
- Claude Haiku 4.572.9%
- GPT-5 nano65.8%
Watch out
Nanonets does not publish when it ran each model. The date shown is when Nanonets last changed that model's result file, 2026-03-23 for every model here; the leaderboard page says "As of April 2026". Nanonets runs each model once per document with the provider's default settings and caps each answer at 8,192 tokens, which can cut off models that think at length. A request that failed, such as an image over a provider's size limit, counts as zero: 6 items for each Claude model shown and 96 for Gemini 3 Pro. Two models on the leaderboard, Gemini 3.1 Pro and GPT-5.4, are not shown: the page's scores for them differ from Nanonets' own result files, so they are held until Nanonets confirms which figures are right. The overall score averages four different tasks, and only one of them is pulling fields out of documents. Nanonets gives different sizes for the test on different pages (about 2,000 documents, 6,406 items, 5,376 items), so no single count is shown.
Reliability not measured by this source.
Nanonets built this benchmark and sells document-processing AI, including Nanonets OCR, which it tests alongside the general models. Only the general models are shown here.
See Nanonets's resultsNot tested here: 22 models
Claude Fable 5, Claude Fable 5.1, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 5, Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.4, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: IDP Core Bench
What's in the test
Invoices, receipts, forms, handwritten pages, charts and scanned text, each sent to the model as an image. 5,376 of the items in Nanonets' result files are PNG images; the format of the rest is not published. The overall score is the average of four tasks: pulling out fields, reading text, reading tables, and answering questions about a document.
- Invoices
- Receipts
- Forms
- Handwritten documents
- Charts
Where results change
By task
The four tasks the overall score averages. Each is scored its own way: fields, text and answers by how close they are to the correct text, tables by how well their rows, columns and cells match.
| Model and setup | Pulling out fields | Reading text | Reading tables | Answering questions about a document |
|---|---|---|---|---|
| Gemini 3 Pro (preview)Sent straight to the model, default settings | 85.7% | 81.8% | 95.8% | 64.1% |
| Claude Sonnet 4.6Sent straight to the model, default settings | 89.5% | 73.7% | 96.3% | 65.2% |
| Claude Opus 4.6Sent straight to the model, default settings | 89.8% | 74.0% | 96.0% | 64.4% |
| Gemini 3 Flash (preview)Sent straight to the model, default settings | 91.1% | 81.7% | 85.6% | 63.5% |
| GPT-5.2Sent straight to the model, default settings | 87.5% | 72.8% | 86.0% | 63.5% |
| GPT-4.1Sent straight to the model, default settings | 87.1% | 75.6% | 73.1% | 63.0% |
| GPT-5 miniSent straight to the model, default settings | 85.7% | 73.0% | 69.5% | 65.0% |
| Claude Haiku 4.5Sent straight to the model, default settings | 85.6% | 65.0% | 81.7% | 59.2% |
| GPT-5 nanoSent straight to the model, default settings | 84.7% | 69.6% | 45.3% | 63.5% |
Technical details for IDP Core Bench
Overall: The overall score averages a model's results across four tasks: pulling key fields from documents, reading printed and handwritten text, reading tables, and answering questions about a document. Each answer is scored by how closely it matches the correct one, so this is not the share answered right.
- Gemini 3 Pro (preview), sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- Claude Sonnet 4.6, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- Claude Opus 4.6, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- Gemini 3 Flash (preview), sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- GPT-5.2, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- GPT-4.1, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- GPT-5 mini, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- Claude Haiku 4.5, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
- GPT-5 nano, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
Source version: IDP Leaderboard page and Nanonets result file for Gemini-3-Pro (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Sonnet 4.6 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Opus 4.6 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Gemini-3-Flash (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5.2 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-4.1 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5-Mini (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Haiku 4.5 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5-Nano (last changed 2026-03-23), downloaded 2026-09-29
Results posted 2026-03-23
Run dates not published by the tester; the dates are when results were posted.
License: MIT, as the IDP Leaderboard site declares ("license":"https://opensource.org/licenses/MIT"); Nanonets agreed by email on 2026-09-28. Checked 2026-09-29.
Sources for this summary:
OmniDocBench
best 92.91 pointsOpenDataLabNewest result 2026-03-31reliability not measuredcaveat
OmniDocBench tests how well a model reads a whole PDF page and turns it into Markdown. It checks the text, the tables and the formulas against a human-checked answer for each page.
Results
All Sent straight to the model · Run date not published; posted 2026-03-31
- Gemini 3 Pro (preview)92.91 points
- Gemini 3 Flash (preview)92.62 points
- GPT-5.286.59 points
Watch out
OmniDocBench does not publish when it ran each model; the date shown is the dated update note on its project page for these models' evaluations. This page shows only the Overall figure; OmniDocBench also publishes separate text, table and formula sub-scores that are not shown here. Data: OmniDocBench, Apache 2.0.
Reliability not measured by this source.
OmniDocBench is run by OpenDataLab.
See OpenDataLab's resultsNot tested here: 28 models
Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.4, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: OmniDocBench
What's in the test
The benchmark is a set of PDF pages covering many document types, layouts and languages. Every page has detailed labels for text, tables, formulas and reading order that people checked by hand.
- Academic literature
- Slides converted from PDF
- Black and white books and textbooks
- Colorful textbooks with images
- Exam papers
- Handwritten notes
- Magazines
- Research and financial reports
- Newspapers
Technical details for OmniDocBench
Overall: OmniDocBench averages three parts on a 0 to 100 scale: text accuracy, table structure and formula accuracy.
- Gemini 3 Pro (preview), sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.
- Gemini 3 Flash (preview), sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.
- GPT-5.2, sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.
Source version: OmniDocBench README at commit 9b46e6da431551535616273c0c57f4510c88b8c5
Results posted 2026-03-31
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-28.
No figures on this page for: Claude Opus 4.5, Claude Opus 5, Claude Sonnet 4.5, Claude Sonnet 5.5, Gemini 2.5 Pro, GPT-5.1, GPT-5.1 Codex, GPT-5.2 Codex, GPT-5.4 Mini.
Also tested by
Their figures are not shown here because permission to reuse them is not yet confirmed.
ExtractBench (Contextual AI)
Contextual AI, checked 2026-09-24A test from Contextual AI of how well AI fills in detailed lists of fields from 35 PDF documents, checked field by field against answers that people validated.
Last updated 2026-02-13
See their resultsMore about this test: ExtractBench (Contextual AI)
What's in the test
35 PDF documents in four areas: finance, academia, hiring and sports. Each comes with a list of fields to fill in, from tens to hundreds of fields, for 12,867 fields in all. One kind of document, a company's regulatory financial filing (an SEC 10-K or 10-Q), asks for 369 fields.
35 documents
- Finance
- Academia
- Hiring
- Sports
Watch out
All 35 documents are in English. The authors say the set is built to show where models fail, not to give precise statistics for smaller groups of documents.
Reliability not measured by this source.
Technical details for ExtractBench (Contextual AI)
Overall Score: This score runs from 0 to 1. Each separate entry in a list counts toward it, and the scoring rules can give a field partial credit.
License: MIT software license; reuse of the paper's scores not stated. Checked 2026-09-24.
What this doesn't tell you
- None of these tests used your documents. The layout, the quality of a scan and the language can all change results.
- Some of the companies that built these tests also sell their own extraction products and test them too. This page shows only the general AI models.
- Results depend on how the model was set up: sent the document in one direct call, or run as a coding agent that can use tools. Compare models only within the same setup.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- DocuBench: Reliability not measured by this source.
- ExtractBench (LlamaIndex): Reliability not measured by this source.
- ParseBench: Reliability not measured by this source.
- IDP Core Bench: Reliability not measured by this source.
- OmniDocBench: Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use