Assurance

Pulling data from documents

Reading invoices, forms, contracts and statements and pulling out the pieces of information you need, such as totals, dates, names and each line of a table. A good result gets every field right. A missing value or a made-up value counts against the model.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Direct sent straight to the modelAgent · Codex coding agentAgent · Claude Code coding agent asked to show where each value came frommeasured measured in these testsauthors report what the test's authors reportcommon practice published advice, not measured hereFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, earned repeat runs, some or all models: Results checked by repeat runs, for some or all models, not yet

Vendor claims have not been collected for any category yet.

DB 91.7%DB 91.7%EB 94.6%EB 94.6%PB 79.85 pointsCX ↗IDP 81.8%IDP 81.8%ODB 92.91 points
ⓘ More about this category
  • Reliability not yet measured
  • 6 sources, 31 models with figures
  • Sources last checked 2026-09-29
  • Ring 4: no rerun evidence yet
  • DB = DocuBench, EB = ExtractBench (LlamaIndex), PB = ParseBench, CX = ExtractBench (Contextual AI), IDP = IDP Core Bench, ODB = OmniDocBench

Who does well

TestBestBest scoreNextNext score
DocuBenchClaude Sonnet 5 Sent straight to the model91.7%GPT-5.5 Sent straight to the model76.5%
ExtractBench (LlamaIndex)GPT-6 Sol Coding agent (Codex), asked to show where each value came fromroughly tied, 3 setupsAlso roughly tied: GPT-5.5 Coding agent (Codex) 93.6%94.6%GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
ParseBenchClaude Opus 5.5 Sent straight to the model, effort highroughly tied79.85 pointsClaude Fable 5.1 Sent straight to the model, default setting78.92 points
IDP Core BenchGemini 3 Pro (preview) Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: Claude Opus 4.6 Sent straight to the model, default settings 81.1%81.8%Claude Sonnet 4.6 Sent straight to the model, default settings81.2%
OmniDocBenchGemini 3 Pro (preview) Sent straight to the modelroughly tied92.91 pointsGemini 3 Flash (preview) Sent straight to the model92.62 points

Where the leader changes

Part of a testLeader thereScore
DocuBench · Photo (JPEG)Gemini 3.5 Flash Sent straight to the model95.2%
DocuBench · HebrewGemini 3.5 Flash Sent straight to the model82.5%
ExtractBench (LlamaIndex) · Short documentsGPT-6 Astra Sent straight to the model97.2%
ExtractBench (LlamaIndex) · Long documentsClaude Opus 4.8 Coding agent (Claude Code)roughly tied, 3 setupsAlso roughly tied: GPT-5.5 Coding agent (Codex), asked to show where each value came from 88.0%; GPT-6 Sol Coding agent (Codex), asked to show where each value came from 87.3%88.1%
ParseBench · TablesClaude Opus 5.5 Sent straight to the model, effort highroughly tiedAlso roughly tied: Claude Opus 5.5 Sent straight to the model, default setting 93.86 points94.25 points
ParseBench · Keeping the text complete and correctClaude Opus 5.5 Sent straight to the model, default settingroughly tied, 5 setupsAlso roughly tied: Claude Opus 5.5 Sent straight to the model, effort high 91.72 points; Claude Opus 5.5 Sent straight to the model, effort low 91.68 points; Claude Fable 5.1 Sent straight to the model, default setting 91.19 points; Gemini 3 Flash (preview) Sent straight to the model, thinking high 90.87 points91.81 points
ParseBench · Formatting that carries meaningGemini 3.5 Flash Sent straight to the model, thinking highroughly tied, 4 setupsAlso roughly tied: GPT-5.6 Sol Sent straight to the model, reasoning high 77.16 points; Claude Opus 5.5 Sent straight to the model, default setting 77.04 points; Claude Opus 5.5 Sent straight to the model, effort high 77.04 points77.85 points
ParseBench · Placing each part on the pageGemini 3.8 Flash Sent straight to the model, thinking highroughly tiedAlso roughly tied: Gemini 3.8 Flash Sent straight to the model, thinking low 72.24 points72.91 points
IDP Core Bench · Pulling out fieldsGemini 3 Flash (preview) Sent straight to the model, default settings91.1%
IDP Core Bench · Reading textGemini 3 Pro (preview) Sent straight to the model, default settingsroughly tiedAlso roughly tied: Gemini 3 Flash (preview) Sent straight to the model, default settings 81.7%81.8%
IDP Core Bench · Reading tablesClaude Sonnet 4.6 Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: Claude Opus 4.6 Sent straight to the model, default settings 96.0%; Gemini 3 Pro (preview) Sent straight to the model, default settings 95.8%96.3%
IDP Core Bench · Answering questions about a documentClaude Sonnet 4.6 Sent straight to the model, default settingsroughly tied, 3 setupsAlso roughly tied: GPT-5 mini Sent straight to the model, default settings 65.0%; Claude Opus 4.6 Sent straight to the model, default settings 64.4%65.2%
Place within each test. Places are never added up across tests.
ModelDocuBenchExtractBench (LlamaIndex)ParseBenchIDP Core BenchOmniDocBench
Claude Haiku 4.524 of 26Sent straight to the model, thinking on8 of 9Sent straight to the model, default settings
Claude Opus 4.623 of 26Sent straight to the model, default setting3 of 9Sent straight to the model, default settings
Claude Opus 4.88 of 12Coding agent (Claude Code)17 of 26Sent straight to the model, default setting
Claude Opus 5.59 of 12Coding agent (Claude Code), asked to show where each value came from1 of 26Sent straight to the model, effort high
Claude Sonnet 51 of 3Sent straight to the model20 of 26Sent straight to the model, default setting
Gemini 3 Flash (preview)4 of 26Sent straight to the model, thinking high4 of 9Sent straight to the model, default settings2 of 3Sent straight to the model
Gemini 3 Pro (preview)1 of 9Sent straight to the model, default settings1 of 3Sent straight to the model
Gemini 3.5 Flash3 of 3Sent straight to the model11 of 12Sent straight to the model5 of 26Sent straight to the model, thinking high
Gemini 3.8 Flash10 of 12Sent straight to the model6 of 26Sent straight to the model, thinking high
GPT-5 mini25 of 26Sent straight to the model, reasoning medium7 of 9Sent straight to the model, default settings
GPT-5.25 of 9Sent straight to the model, default settings3 of 3Sent straight to the model
GPT-5.4 Nano12 of 12Sent straight to the model26 of 26Sent straight to the model, default setting
GPT-5.52 of 3Sent straight to the model3 of 12Coding agent (Codex)15 of 26Sent straight to the model, reasoning medium
GPT-5.6 Luna6 of 12Coding agent (Codex), asked to show where each value came from14 of 26Sent straight to the model, reasoning max
GPT-5.6 Sol2 of 12Coding agent (Codex), asked to show where each value came from3 of 26Sent straight to the model, reasoning high
GPT-5.6 Terra4 of 12Coding agent (Codex), asked to show where each value came from13 of 26Sent straight to the model, reasoning medium
GPT-6 Astra5 of 12Sent straight to the model9 of 26Sent straight to the model, reasoning low
GPT-6 Luna7 of 12Coding agent (Codex), asked to show where each value came from16 of 26Sent straight to the model, reasoning max
GPT-6 Sol1 of 12Coding agent (Codex), asked to show where each value came from11 of 26Sent straight to the model, reasoning high

The tests here put 34 pairs of models in a different order.

Show all 34 pairs placed in a different order
  • In ParseBench, Claude Haiku 4.5, sent straight to the model, thinking on, placed above GPT-5 mini, sent straight to the model, reasoning medium; in IDP Core Bench, GPT-5 mini, sent straight to the model, default settings, placed above Claude Haiku 4.5, sent straight to the model, default settings.
  • In ParseBench, Gemini 3 Flash (preview), sent straight to the model, thinking high, placed above Claude Opus 4.6, sent straight to the model, default setting; in IDP Core Bench, Claude Opus 4.6, sent straight to the model, default settings, placed above Gemini 3 Flash (preview), sent straight to the model, default settings.
  • In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, default setting.
  • In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Claude Opus 4.8, sent straight to the model, default setting.
  • In ExtractBench (LlamaIndex), Claude Opus 4.8, coding agent (Claude Code), placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above Claude Opus 4.8, sent straight to the model, default setting.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, reasoning high.
  • In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
  • In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from; in ParseBench, Claude Opus 5.5, sent straight to the model, effort high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
  • In DocuBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Claude Sonnet 5, sent straight to the model, default setting.
  • In DocuBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.5, sent straight to the model; in ParseBench, GPT-5.5, sent straight to the model, reasoning medium, placed above Claude Sonnet 5, sent straight to the model, default setting.
  • In ExtractBench (LlamaIndex), Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above Gemini 3.8 Flash, sent straight to the model, thinking high.
  • In DocuBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
  • In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.5 Flash, sent straight to the model; in ParseBench, Gemini 3.5 Flash, sent straight to the model, thinking high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-6 Astra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Astra, sent straight to the model, reasoning low.
  • In ExtractBench (LlamaIndex), GPT-6 Luna, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Luna, sent straight to the model, reasoning max.
  • In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above Gemini 3.8 Flash, sent straight to the model; in ParseBench, Gemini 3.8 Flash, sent straight to the model, thinking high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Luna, sent straight to the model, reasoning max, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Terra, sent straight to the model, reasoning medium, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-5.5, coding agent (Codex), placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-5.5, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from; in ParseBench, GPT-5.6 Sol, sent straight to the model, reasoning high, placed above GPT-6 Sol, sent straight to the model, reasoning high.
  • In ExtractBench (LlamaIndex), GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from, placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-5.6 Terra, sent straight to the model, reasoning medium.
  • In ExtractBench (LlamaIndex), GPT-6 Sol, coding agent (Codex), asked to show where each value came from, placed above GPT-6 Astra, sent straight to the model; in ParseBench, GPT-6 Astra, sent straight to the model, reasoning low, placed above GPT-6 Sol, sent straight to the model, reasoning high.

By input type

Each test is listed on its own.

DocuBench

58 documents

Most accurate

  1. Claude Sonnet 5 Sent straight to the model92.3%
  2. GPT-5.5 Sent straight to the model72.5%
  3. Gemini 3.5 Flash Sent straight to the model67.0%

No cost published

ExtractBench (LlamaIndex)

370 documents

Most accurate

  1. GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
  2. GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
  3. GPT-5.5 Coding agent (Codex)93.6%
  4. GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
  5. GPT-6 Astra Sent straight to the model91.9%
  6. GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
  7. GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
  8. Claude Opus 4.8 Coding agent (Claude Code)87.1%

What this test found for PDF

  • In ExtractBench (LlamaIndex), every model and setup scored lower on long documents than on short ones, by 2.0 to 65.5 points (14 compared). See the evidence
  • The authors of ExtractBench (LlamaIndex) report that when a whole long document goes to a model in one request, the model often cuts long lists short. See the evidence
  • On long documents in ExtractBench (LlamaIndex), the best agent setup scored 88.1% and the best single request scored 35.8%, a gap of 52.3 points. See the evidence
  • On short documents in ExtractBench (LlamaIndex), asking the model to show where each value came from changed accuracy by at most 0.5 points (2 models tried both ways). See the evidence
  • On long documents it was different: asking for sources changed accuracy by as much as 10.1 points, for Claude Opus 4.8, coding agent (Claude Code); and GPT-5.5, coding agent (Codex). See the evidence

Cost: see Quality and cost

OmniDocBench

PDF pages sent to the model as images

Most accurate

  1. Gemini 3 Pro (preview) Sent straight to the model92.91 points
  2. Gemini 3 Flash (preview) Sent straight to the model92.62 points
  3. GPT-5.2 Sent straight to the model86.59 points

No cost published

Quality and cost

In ExtractBench (LlamaIndex) only. Cost as published by LlamaIndex. Your costs will differ.

Most accurate, and the cheapest within 1 point of the best: GPT-6 Sol Coding agent (Codex), asked to show where each value came from 94.6% · $0.098 per page

Best at each price

  1. 1GPT-5.4 Nano Sent straight to the model 74.9% · $0.0021 per page
  2. 2Gemini 3.8 Flash Sent straight to the model 80.7% · $0.0043 per page
  3. 3GPT-6 Luna Coding agent (Codex), asked to show where each value came from 91.0% · $0.0097 per page
  4. 4GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from 91.2% · $0.012 per page
  5. 5GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from 92.3% · $0.0945 per page
  6. 6GPT-6 Sol Coding agent (Codex), asked to show where each value came from 94.6% · $0.098 per pagebest quality

In ParseBench only. Cost as published by LlamaIndex. Your costs will differ.

Most accurate, and the cheapest within 1 point of the best: Claude Opus 5.5 Sent straight to the model, effort high 79.85 points · $0.0613 per page

Best at each price

  1. 1GPT-6 Luna Sent straight to the model, reasoning none 52.93 points · $0.0009 per page
  2. 2GPT-6 Luna Sent straight to the model, reasoning low 59.32 points · $0.0011 per page
  3. 3GPT-6 Luna Sent straight to the model, reasoning medium 62.38 points · $0.0014 per page
  4. 4GPT-6 Luna Sent straight to the model, reasoning high 64.21 points · $0.0021 per page
  5. 5GPT-6 Luna Sent straight to the model, reasoning xhigh 64.32 points · $0.0029 per page
  6. 6GPT-5.6 Luna Sent straight to the model, reasoning high 67.37 points · $0.0059 per page
  7. 7Gemini 3 Flash (preview) Sent straight to the model, thinking minimal 71.04 points · $0.0065 per page
  8. 8Gemini 3.7 Flash Sent straight to the model, thinking high 71.27 points · $0.0191 per page
  9. 9Gemini 3 Flash (preview) Sent straight to the model, thinking medium 75 points · $0.0198 per page
  10. 10Gemini 3 Flash (preview) Sent straight to the model, thinking high 75.05 points · $0.0241 per page
  11. 11Claude Opus 5.5 Sent straight to the model, default setting 78.01 points · $0.0579 per page
  12. 12Claude Opus 5.5 Sent straight to the model, effort high 79.85 points · $0.0613 per pagebest quality

Where it goes wrong

What helps

Next steps for a small business

  1. Test on a sample of your own documents and compare the answers with ones you have already checked.P1 behind step 1: Test on a sample of your own documents and compare the answers with ones you have already checked.
  2. For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.F5 behind step 2: For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.F2 behind step 2: For long statements or registers, use an agent setup rather than one request, and weigh its higher cost. In the tests we show, agent setups led on long documents.
  3. If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.F6 behind step 3: If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.F7 behind step 3: If you want the model to show where each value came from, test it both ways on your own files first. In the tests we show, it made little difference on short documents and a large difference on long ones.
  4. Keep a person reviewing anything you will pay, file or sign.P2 behind step 4: Keep a person reviewing anything you will pay, file or sign.
  5. If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.F3 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.F4 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.P3 behind step 5: If your documents are images, such as photos or scans, or are not in English, test those separately. The leader changed on image files and on Hebrew documents in DocuBench.

F: a finding above. P: a common practice above.

Sample jobs

Everyday jobs, each matched to the test that is most like it. This is simple arithmetic on the test's score. Your documents will differ.

Invoices for accounts payable

Pulling 12 fields, such as vendor, invoice number, date and total, from each of 500 invoices a month: 6,000 fields. · DocuBench
best: about 1 in 12 wronglowest: about 1 in 4

Pulling 12 fields, such as vendor, invoice number, date and total, from each of 500 invoices a month: 6,000 fields.

Closest test: DocuBench. DocuBench scores each field as right or wrong, and its documents include invoices, statements and bills like the ones an office handles.

  • Claude Sonnet 5, sent straight to the model (91.7%): about 1 in 12 fields wrong, roughly 500 of 6,000 a month.
  • Gemini 3.5 Flash, sent straight to the model (73.0%): about 1 in 4 fields wrong, roughly 1,600 of 6,000 a month.

Phone photos of receipts

Pulling 10 fields, such as store, date and total, from each of 100 receipts photographed with a phone: 1,000 fields a month. · DocuBench, Photo (JPEG)
best: about 1 in 21 wronglowest: about 1 in 4

Pulling 10 fields, such as store, date and total, from each of 100 receipts photographed with a phone: 1,000 fields a month.

Closest test: DocuBench, Photo (JPEG) only. This DocuBench group covers five photo (JPEG) files, including one receipt photo. It is a small, mixed group, so it is only a loose match for phone photos of receipts.

  • Gemini 3.5 Flash, sent straight to the model (95.2%): about 1 in 21 fields wrong, roughly 48 of 1,000 a month.
  • Claude Sonnet 5, sent straight to the model (75.9%): about 1 in 4 fields wrong, roughly 240 of 1,000 a month.

Long statements and registers

Pulling every line from long documents, such as a year of account statements or a check register that runs past 50 pages, where missing even a few lines matters. · ExtractBench (LlamaIndex), Long documents
Open for the figures

Pulling every line from long documents, such as a year of account statements or a check register that runs past 50 pages, where missing even a few lines matters.

Closest test: ExtractBench (LlamaIndex), Long documents only. The long documents in ExtractBench (LlamaIndex) each run more than 50 pages, like its example check register with 3,308 payment lines.

  • Claude Opus 4.8, coding agent (Claude Code): 88.1%
  • GPT-5.5, coding agent (Codex), asked to show where each value came from: 88.0%
  • GPT-6 Sol, coding agent (Codex), asked to show where each value came from: 87.3%
  • GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from: 82.7%
  • GPT-6 Luna, coding agent (Codex), asked to show where each value came from: 80.9%
  • Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from: 80.8%
  • GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from: 80.0%
  • GPT-5.5, coding agent (Codex): 78.9%
  • Claude Opus 4.8, coding agent (Claude Code), asked to show where each value came from: 77.9%
  • GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from: 77.4%
  • GPT-5.4 Nano, sent straight to the model: 35.8%
  • GPT-6 Astra, sent straight to the model: 31.7%
  • Gemini 3.8 Flash, sent straight to the model: 28.7%
  • Gemini 3.5 Flash, sent straight to the model: 27.9%

This score is not a simple share of correct answers, so we do not turn it into a count.

The tests behind this

DocuBench

best 91.7%DocuPipe72 documentsNewest result 2026-07-17reliability not measuredcaveat

A test of how well AI pulls requested information out of 72 hard documents, checked against answers that people verified by hand.

Results

All Sent straight to the model · Run date not published; posted 2026-07-17

  • Claude Sonnet 591.7%
  • GPT-5.576.5%
  • Gemini 3.5 Flash73.0%

Watch out

Most of the test is in English: 47 of the 72 documents. Several languages have only one to three documents each, so treat their results as examples. The authors say 72 documents are enough for careful inspection but not for broad claims about every kind of document.

Reliability not measured by this source.

DocuPipe built this benchmark and sells a document extraction product, which it tests alongside the general models. Only the general models are shown here.

See DocuPipe's results
Not tested here: 28 models

Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.2, GPT-5.4, GPT-5.4 Nano, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: DocuBench

What's in the test

Each test item is one document, a list of the fields to pull out of it, and the correct answer for each field. The documents come in 10 file types, from PDFs and photos to spreadsheets and web pages, and in 12 languages. Most come from public sources such as government publications and sample documents that companies published; a few were created by the benchmark's authors.

72 documents

  • Invoices
  • Statements
  • Utility bills
  • Annual reports
  • Payslips
  • Purchase orders
  • Shipping waybills
  • Healthcare forms
  • Engineering drawings
  • Insurance declarations
  • Dictionaries
  • Directories
  • Auction catalogs
  • Government registers
  • Spreadsheets

An example

CPS Energy electric bill

A two-page home electric bill that the utility published as a sample, with personal details blacked out. The model must pull out the billing period, the meter readings, each charge line and the total, and must not count the total row as a charge.

Billing period:
September 18, 2021 to October 18, 2021
Current meter reading:
48,374
Electricity used:
463 kilowatt-hours
One charge line:
Service Availability Charge, $8.75
Total electric bill:
$58.35
Open this document

Where results change

By kind of file

The same test split by the kind of file the document came in.

Model and setupPhoto (JPEG)PDF
Claude Sonnet 5Sent straight to the model75.9%92.3%
GPT-5.5Sent straight to the model88.3%72.5%
Gemini 3.5 FlashSent straight to the model95.2%67.0%
By language

The same test split by the document's language.

Model and setupEnglishHebrew
Claude Sonnet 5Sent straight to the model95.2%70.6%
GPT-5.5Sent straight to the model80.2%76.5%
Gemini 3.5 FlashSent straight to the model75.1%82.5%
Technical details for DocuBench

Macro-average field accuracy: The average share of fields pulled out correctly per document. A failed request counts every field on that document as missed.

  • Claude Sonnet 5, sent straight to the model: Direct API call, one request per document. Tools: None. Input: PDF sent as a document. Full output budget, per the source's README.
  • GPT-5.5, sent straight to the model: Direct API call, one request per document. Tools: None. Input: File upload. Full output budget, per the source's README.
  • Gemini 3.5 Flash, sent straight to the model: Direct API call, one request per document. Tools: None. Input: File sent inline. Full output budget, per the source's README.

Source version: DocuBench 0.2.0

Results posted 2026-07-17

Run dates not published by the tester; the dates are when results were posted.

License: CC BY 4.0 (results folder). Checked 2026-09-25.

Based on one or a few documents; treat as an example, not a pattern.

  • By kind of file, CSV file (1 document): Claude Sonnet 5, sent straight to the model: 97.1%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, Word document (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 83.3%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, Web page (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, Image (PNG) (1 document): Claude Sonnet 5, sent straight to the model: 95.7%; GPT-5.5, sent straight to the model: 79.4%; Gemini 3.5 Flash, sent straight to the model: 92.2%
  • By kind of file, Scan (TIFF) (1 document): Claude Sonnet 5, sent straight to the model: 85.7%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, Text file (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, Spreadsheet (Excel) (1 document): Claude Sonnet 5, sent straight to the model: 93.2%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By kind of file, XML file (2 documents): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By language, Arabic (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By language, German (2 documents): Claude Sonnet 5, sent straight to the model: 96.1%; GPT-5.5, sent straight to the model: 96.2%; Gemini 3.5 Flash, sent straight to the model: 99.6%
  • By language, Spanish (3 documents): Claude Sonnet 5, sent straight to the model: 98.3%; GPT-5.5, sent straight to the model: 60.7%; Gemini 3.5 Flash, sent straight to the model: 35.8%
  • By language, French (1 document): Claude Sonnet 5, sent straight to the model: 81.0%; GPT-5.5, sent straight to the model: 100.0%; Gemini 3.5 Flash, sent straight to the model: 94.8%
  • By language, Hindi (2 documents): Claude Sonnet 5, sent straight to the model: 85.8%; GPT-5.5, sent straight to the model: 8.5%; Gemini 3.5 Flash, sent straight to the model: 8.2%
  • By language, Italian (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 96.5%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By language, Japanese (4 documents): Claude Sonnet 5, sent straight to the model: 79.7%; GPT-5.5, sent straight to the model: 48.1%; Gemini 3.5 Flash, sent straight to the model: 52.9%
  • By language, Dutch (1 document): Claude Sonnet 5, sent straight to the model: 100.0%; GPT-5.5, sent straight to the model: 76.7%; Gemini 3.5 Flash, sent straight to the model: 100.0%
  • By language, Portuguese (1 document): Claude Sonnet 5, sent straight to the model: 96.4%; GPT-5.5, sent straight to the model: 97.6%; Gemini 3.5 Flash, sent straight to the model: 96.4%
  • By language, Chinese (2 documents): Claude Sonnet 5, sent straight to the model: 90.2%; GPT-5.5, sent straight to the model: 72.5%; Gemini 3.5 Flash, sent straight to the model: 59.7%

ExtractBench (LlamaIndex)

best 94.6%LlamaIndex370 documentsNewest result 2026-09-23reliability not measuredcaveat

A test of how well AI fills in a set list of fields from 370 business documents, from short forms to reports of more than 50 pages, checked field by field against the correct answers.

Results

Run date not published; posted 2026-09-23

Each model's best setting in this test

  • GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
  • GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
  • GPT-5.5 Coding agent (Codex)93.6%
  • GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
  • GPT-6 Astra Sent straight to the model91.9%
  • GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
  • GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
  • Claude Opus 4.8 Coding agent (Claude Code)87.1%
Show all 14 results (6 not shown above) from ExtractBench (LlamaIndex)
  • GPT-6 Sol Coding agent (Codex), asked to show where each value came from94.6%
  • GPT-5.6 Sol Coding agent (Codex), asked to show where each value came from93.8%
  • GPT-5.5 Coding agent (Codex)93.6%
  • GPT-5.5 Coding agent (Codex), asked to show where each value came from93.3%
  • GPT-5.6 Terra Coding agent (Codex), asked to show where each value came from92.3%
  • GPT-6 Astra Sent straight to the model91.9%
  • GPT-5.6 Luna Coding agent (Codex), asked to show where each value came from91.2%
  • GPT-6 Luna Coding agent (Codex), asked to show where each value came from91.0%
  • Claude Opus 4.8 Coding agent (Claude Code)87.1%
  • Claude Opus 4.8 Coding agent (Claude Code), asked to show where each value came from86.8%
  • Claude Opus 5.5 Coding agent (Claude Code), asked to show where each value came from86.1%
  • Gemini 3.8 Flash Sent straight to the model80.7%
  • Gemini 3.5 Flash Sent straight to the model79.8%
  • GPT-5.4 Nano Sent straight to the model74.9%

Watch out

The leaderboard does not say when each model was run; the date shown is when the leaderboard was last updated. The authors report that models sent a whole document in one direct call do well on short documents but often cut long lists short, while coding agents stay more accurate at much higher cost.

Reliability not measured by this source.

LlamaIndex built this benchmark and sells a document extraction product, LlamaExtract, which it tests alongside the general models. Only the general models are shown here.

See LlamaIndex's results
Not tested here: 19 models

Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.2, GPT-5.4.

More about this test: ExtractBench (LlamaIndex)

What's in the test

The documents come from public records such as company and regulatory filings, government purchasing and customs forms, court exhibits and Texas energy filings. Of the 370, 45 are long lists created by the benchmark's authors from real layouts. The test sorts them by length: 252 short (10 pages or fewer), 98 medium (11 to 50 pages) and 20 long (more than 50 pages).

370 documents

  • Finance and fund holdings
  • Energy regulatory forms
  • Government purchasing and customs forms
  • Auto valuation
  • Supply chain
  • Healthcare payment notices
  • Legal and bankruptcy filings
  • Real estate

An example

Freer ISD check register

A real report from the Freer Independent School District in Texas listing every check the district wrote in one school year. It runs 123 pages. The model must list every payment line and copy each check's number, date and payee onto the lines below it, where they are not printed again.

District name:
Freer ISD
Report title:
YTD Check Register
Report period:
September 1, 2022 to August 31, 2023
Total amount:
$5,991,776.11
Payment lines to find:
3,308
Open this document

Where results change

By document length

The same test split by how long the documents are.

Model and setupShort documentsMedium documentsLong documents
GPT-6 SolCoding agent (Codex), asked to show where each value came from96.0%92.4%87.3%
GPT-5.6 SolCoding agent (Codex), asked to show where each value came from96.0%90.2%82.7%
GPT-5.5Coding agent (Codex)95.7%91.2%78.9%
GPT-5.5Coding agent (Codex), asked to show where each value came from95.6%88.7%88.0%
GPT-5.6 TerraCoding agent (Codex), asked to show where each value came from95.6%86.7%77.4%
GPT-6 AstraSent straight to the model97.2%90.6%31.7%
GPT-5.6 LunaCoding agent (Codex), asked to show where each value came from93.5%87.4%80.0%
GPT-6 LunaCoding agent (Codex), asked to show where each value came from93.1%87.7%80.9%
Claude Opus 4.8Coding agent (Claude Code)90.1%79.2%88.1%
Claude Opus 4.8Coding agent (Claude Code), asked to show where each value came from90.5%79.2%77.9%
Claude Opus 5.5Coding agent (Claude Code), asked to show where each value came from89.6%78.0%80.8%
Gemini 3.8 FlashSent straight to the model88.1%72.2%28.7%
Gemini 3.5 FlashSent straight to the model87.9%69.8%27.9%
GPT-5.4 NanoSent straight to the model77.4%76.4%35.8%
Technical details for ExtractBench (LlamaIndex)

Unified value F1: An F1 score for extracted values, counting every item in a list. It rewards values that match the correct answer, allowing for small formatting differences such as how dates are written. It penalizes both wrong and missing values, and it is averaged evenly across documents. Rows marked Evidence are runs where the agent was also asked to cite the page and location behind each value it extracted.

  • GPT-6 Sol, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • GPT-5.6 Sol, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • GPT-5.5, coding agent (Codex): Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: standard.
  • GPT-5.5, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • GPT-5.6 Terra, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
  • GPT-5.6 Luna, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • GPT-6 Luna, coding agent (Codex), asked to show where each value came from: Codex coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • Claude Opus 4.8, coding agent (Claude Code): Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: standard.
  • Claude Opus 4.8, coding agent (Claude Code), asked to show where each value came from: Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • Claude Opus 5.5, coding agent (Claude Code), asked to show where each value came from: Claude Code coding agent. Tools: The agent's own tools. Input: Documents as provided by the benchmark. Variant: Evidence.
  • Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
  • Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.
  • GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: Documents as provided by the benchmark. Variant: standard.

Source version: ExtractBench leaderboard, 2026-09-23

Results posted 2026-09-23

Run dates not published by the tester; the dates are when results were posted.

License: Apache-2.0; LlamaIndex confirmed public use of the results by email on 2026-09-24. Checked 2026-09-26.

ParseBench

best 79.85 pointsLlamaIndex2,078 pagesNewest result 2026-09-25reliability not measuredcaveat

A test from LlamaIndex of how well AI turns document pages into clean, structured text: tables, charts, the words on the page, the formatting that carries meaning, and where each part sits on the page.

Results

Each model's best setting in this test

  • Claude Opus 5.5 Sent straight to the model, effort high, Run date not published; posted 2026-09-2479.85 points
  • Claude Fable 5.1 Sent straight to the model, default setting, Run date not published; posted 2026-09-0278.92 points
  • GPT-5.6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2575.35 points
  • Gemini 3 Flash (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-04-2175.05 points
  • Gemini 3.5 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2574.02 points
  • Gemini 3.8 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-0472.08 points
  • Gemini 3.7 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2571.27 points
  • Claude Fable 5 Sent straight to the model, default setting, Run date not published; posted 2026-06-1070.78 points
Show all 61 results (53 not shown above) from ParseBench
  • Claude Opus 5.5 Sent straight to the model, effort high, Run date not published; posted 2026-09-2479.85 points
  • Claude Fable 5.1 Sent straight to the model, default setting, Run date not published; posted 2026-09-0278.92 points
  • Claude Opus 5.5 Sent straight to the model, default setting, Run date not published; posted 2026-09-2278.01 points
  • GPT-5.6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2575.35 points
  • Gemini 3 Flash (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-04-2175.05 points
  • Gemini 3 Flash (preview) Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2575 points
  • Claude Opus 5.5 Sent straight to the model, effort low, Run date not published; posted 2026-09-2474.47 points
  • Gemini 3.5 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2574.02 points
  • Gemini 3.8 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-0472.08 points
  • Gemini 3.7 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2571.27 points
  • Gemini 3 Flash (preview) Sent straight to the model, thinking minimal, Run date not published; posted 2026-04-2171.04 points
  • Claude Fable 5 Sent straight to the model, default setting, Run date not published; posted 2026-06-1070.78 points
  • Gemini 3.7 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2570.78 points
  • Gemini 3.8 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2570.7 points
  • GPT-6 Astra Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2470.67 points
  • Gemini 3.8 Flash Sent straight to the model, thinking low, Run date not published; posted 2026-09-0470.15 points
  • Gemini 3.6 Flash Sent straight to the model, thinking high, Run date not published; posted 2026-09-2570.03 points
  • Gemini 3.5 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-06-0169.92 points
  • GPT-6 Sol Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2569.69 points
  • Gemini 3.7 Flash Sent straight to the model, thinking low, Run date not published; posted 2026-09-2569.57 points
  • Gemini 3.1 Pro (preview) Sent straight to the model, default setting, Run date not published; posted 2026-04-2169.14 points
  • GPT-5.6 Terra Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2568.39 points
  • GPT-5.6 Luna Sent straight to the model, reasoning max, Run date not published; posted 2026-09-2468.34 points
  • GPT-5.6 Luna Sent straight to the model, reasoning xhigh, Run date not published; posted 2026-09-2468.24 points
  • GPT-6 Sol Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2268.19 points
  • GPT-5.6 Terra Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2568.07 points
  • GPT-5.5 Sent straight to the model, reasoning medium, Run date not published; posted 2026-04-2467.76 points
  • GPT-5.6 Luna Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2467.37 points
  • Gemini 3.6 Flash Sent straight to the model, thinking medium, Run date not published; posted 2026-08-0466.78 points
  • GPT-6 Sol Sent straight to the model, reasoning none, Run date not published; posted 2026-09-2266.1 points
  • GPT-6 Luna Sent straight to the model, reasoning max, Run date not published; posted 2026-09-2465.77 points
  • Gemini 3.6 Flash Sent straight to the model, thinking minimal, Run date not published; posted 2026-08-0465.6 points
  • GPT-5.5 Sent straight to the model, reasoning none, Run date not published; posted 2026-04-2464.39 points
  • GPT-6 Luna Sent straight to the model, reasoning xhigh, Run date not published; posted 2026-09-2464.32 points
  • GPT-6 Luna Sent straight to the model, reasoning high, Run date not published; posted 2026-09-2464.21 points
  • GPT-5.6 Terra Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1364.17 points
  • Claude Opus 4.8 Sent straight to the model, default setting, Run date not published; posted 2026-06-0163.7 points
  • GPT-5.6 Luna Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2463.58 points
  • Claude Opus 4.7 Sent straight to the model, default setting, Run date not published; posted 2026-04-2163.34 points
  • Gemini 3.5 Flash Sent straight to the model, thinking minimal, Run date not published; posted 2026-06-0163.09 points
  • GPT-6 Luna Sent straight to the model, reasoning medium, Run date not published; posted 2026-09-2262.38 points
  • GPT-5.4 Sent straight to the model, reasoning none, Run date not published; posted 2026-04-2162.23 points
  • Claude Sonnet 5 Sent straight to the model, default setting, Run date not published; posted 2026-07-0262.13 points
  • GPT-5.6 Sol Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1362.12 points
  • Gemini 3.1 Pro (preview) Sent straight to the model, thinking low, Run date not published; posted 2026-09-2560.82 points
  • Gemini 3.5 Flash-Lite Sent straight to the model, thinking high, Run date not published; posted 2026-09-2560.64 points
  • GPT-5.6 Luna Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2460.18 points
  • GPT-6 Luna Sent straight to the model, reasoning low, Run date not published; posted 2026-09-2459.32 points
  • Gemini 3.1 Flash-Lite (preview) Sent straight to the model, default setting, Run date not published; posted 2026-04-2158.32 points
  • Gemini 3.1 Flash-Lite (preview) Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2558.13 points
  • Gemini 3.1 Flash-Lite (preview) Sent straight to the model, thinking high, Run date not published; posted 2026-09-2557.29 points
  • Gemini 3.5 Flash-Lite Sent straight to the model, default setting, Run date not published; posted 2026-08-0457.15 points
  • GPT-5.6 Luna Sent straight to the model, reasoning none, Run date not published; posted 2026-07-1356.32 points
  • Gemini 3.5 Flash-Lite Sent straight to the model, thinking medium, Run date not published; posted 2026-09-2554.3 points
  • Claude Opus 4.6 Sent straight to the model, default setting, Run date not published; posted 2026-04-2154.07 points
  • Claude Haiku 4.5 Sent straight to the model, thinking on, Run date not published; posted 2026-04-2153.12 points
  • GPT-6 Luna Sent straight to the model, reasoning none, Run date not published; posted 2026-09-2252.93 points
  • GPT-5 mini Sent straight to the model, reasoning medium, Run date not published; posted 2026-04-2151.52 points
  • GPT-5 mini Sent straight to the model, reasoning minimal, Run date not published; posted 2026-04-2146.83 points
  • Claude Haiku 4.5 Sent straight to the model, thinking off, Run date not published; posted 2026-04-2145.17 points
  • GPT-5.4 Nano Sent straight to the model, default setting, Run date not published; posted 2026-04-2143.35 points

Watch out

LlamaIndex does not publish when it ran each model. The date shown is the last time LlamaIndex changed that model's row in its results file, which can be later than the run, for example when it reprices a row. Each setting LlamaIndex tried is its own row, so one model can appear up to six times. LlamaIndex works out each cost from the tokens that run used, priced at the model maker's list price when the row was last priced; when prices change it reprices the row without running it again, so a cost can change while the score stays the same. Each model is run once, one page at a time; LlamaIndex publishes no repeat runs. An Anthropic employee contributed three changes to ParseBench's code on 2026-08-04: two to how formatting is scored and one option that none of the leaderboard's runs use. Every change to every row shown was made by LlamaIndex.

Reliability not measured by this source.

LlamaIndex built this benchmark and sells a document parsing product, LlamaParse, which it tests alongside the general models. LlamaParse held the top place on the leaderboard when we checked on 2026-09-29. Only the general models are shown here.

See LlamaIndex's results
Not tested here: 5 models

Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-4.1, GPT-5 nano, GPT-5.2.

More about this test: ParseBench

What's in the test

About 2,000 pages checked by people, from more than 1,200 public documents in insurance, finance, government and other fields. Most pages are PDF; 42, all in the page-placement part, are JPG or PNG images, which are sent to the model as images. ParseBench publishes no separate scores for them.

2,078 pages

  • Tables
  • Charts
  • Text pages
  • Page layout

Where results change

By what is checked

The five scores the overall averages, each from 0 to 100 and each scored its own way.

Model and setupTablesChartsKeeping the text complete and correctFormatting that carries meaningPlacing each part on the page
Claude Opus 5.5Sent straight to the model, effort high94.25 points70.89 points91.72 points77.04 points65.33 points
Claude Fable 5.1Sent straight to the model, default setting91.52 points67.06 points91.19 points76.52 points68.3 points
Claude Opus 5.5Sent straight to the model, default setting93.86 points64.12 points91.81 points77.04 points63.2 points
GPT-5.6 SolSent straight to the model, reasoning high91.39 points68.68 points88.16 points77.16 points51.35 points
Gemini 3 Flash (preview)Sent straight to the model, thinking high91.5 points64.79 points90.87 points68.31 points59.77 points
Gemini 3 Flash (preview)Sent straight to the model, thinking medium91.01 points61.56 points88.67 points67.95 points65.79 points
Claude Opus 5.5Sent straight to the model, effort low92.69 points53.52 points91.68 points75.21 points59.27 points
Gemini 3.5 FlashSent straight to the model, thinking high90.44 points45.27 points88.92 points77.85 points67.6 points
Gemini 3.8 FlashSent straight to the model, thinking high89.12 points32.34 points89.68 points76.36 points72.91 points
Gemini 3.7 FlashSent straight to the model, thinking high90.16 points32.07 points88.6 points74.98 points70.56 points
Gemini 3 Flash (preview)Sent straight to the model, thinking minimal89.85 points64.83 points86.19 points58.35 points55.97 points
Claude Fable 5Sent straight to the model, default setting89.79 points52.21 points90.02 points72.62 points49.24 points
Gemini 3.7 FlashSent straight to the model, thinking medium88.66 points35.11 points88.45 points73.05 points68.63 points
Gemini 3.8 FlashSent straight to the model, thinking medium89.08 points32.72 points88.42 points71.8 points71.5 points
GPT-6 AstraSent straight to the model, reasoning low93.17 points35.08 points89.6 points75.52 points59.99 points
Gemini 3.8 FlashSent straight to the model, thinking low88.18 points35.09 points88.25 points66.99 points72.24 points
Gemini 3.6 FlashSent straight to the model, thinking high89.3 points34.53 points88.21 points73.11 points64.98 points
Gemini 3.5 FlashSent straight to the model, thinking medium91.12 points44.01 points90.19 points59.14 points65.14 points
GPT-6 SolSent straight to the model, reasoning high92.63 points48.05 points87.81 points67.91 points52.06 points
Gemini 3.7 FlashSent straight to the model, thinking low87.96 points31.84 points88.44 points68.89 points70.71 points
Gemini 3.1 Pro (preview)Sent straight to the model, default setting91 points41.13 points90.16 points52.43 points70.99 points
GPT-5.6 TerraSent straight to the model, reasoning medium86.72 points67.86 points87.06 points68.47 points31.82 points
GPT-5.6 LunaSent straight to the model, reasoning max88.27 points49.16 points86.57 points64.95 points52.73 points
GPT-5.6 LunaSent straight to the model, reasoning xhigh89.42 points49.44 points87.13 points66.51 points48.69 points
GPT-6 SolSent straight to the model, reasoning medium90.27 points46.46 points87.3 points67.88 points49.05 points
GPT-5.6 TerraSent straight to the model, reasoning low86.91 points67.71 points86.83 points67.92 points30.97 points
GPT-5.5Sent straight to the model, reasoning medium90.05 points65.53 points86.81 points60.12 points36.28 points
GPT-5.6 LunaSent straight to the model, reasoning high88.75 points49.25 points86.99 points67.14 points44.74 points
Gemini 3.6 FlashSent straight to the model, thinking medium89.47 points31 points87.7 points58.46 points67.27 points
GPT-6 SolSent straight to the model, reasoning none89.49 points42.02 points86.38 points67.32 points45.27 points
GPT-6 LunaSent straight to the model, reasoning max87.42 points45.71 points84.44 points61.81 points49.48 points
Gemini 3.6 FlashSent straight to the model, thinking minimal85.5 points24.76 points83.3 points64.49 points69.94 points
GPT-5.5Sent straight to the model, reasoning none89.31 points59.11 points87.17 points64.46 points21.9 points
GPT-6 LunaSent straight to the model, reasoning xhigh88.01 points42.38 points85 points65.34 points40.86 points
GPT-6 LunaSent straight to the model, reasoning high87.97 points44.28 points85.27 points63.58 points39.97 points
GPT-5.6 TerraSent straight to the model, reasoning none86.23 points63.75 points82.56 points59.96 points28.35 points
Claude Opus 4.8Sent straight to the model, default setting89.65 points49.75 points89.02 points71.38 points18.69 points
GPT-5.6 LunaSent straight to the model, reasoning medium85.32 points46.91 points85.27 points65.8 points34.6 points
Claude Opus 4.7Sent straight to the model, default setting87.17 points55.84 points90.26 points69.42 points13.99 points
Gemini 3.5 FlashSent straight to the model, thinking minimal86.24 points18.74 points88.22 points68.09 points54.18 points
GPT-6 LunaSent straight to the model, reasoning medium86.74 points42.97 points84.94 points61.13 points36.14 points
GPT-5.4Sent straight to the model, reasoning none83.89 points65.22 points85.57 points59.52 points16.95 points
Claude Sonnet 5Sent straight to the model, default setting86.65 points60.53 points86.51 points64.93 points12.03 points
GPT-5.6 SolSent straight to the model, reasoning none89.34 points59.57 points86.17 points50.42 points25.09 points
Gemini 3.1 Pro (preview)Sent straight to the model, thinking low90.29 points11.43 points87.11 points45.56 points69.7 points
Gemini 3.5 Flash-LiteSent straight to the model, thinking high83.31 points19.27 points86.85 points54.03 points59.75 points
GPT-5.6 LunaSent straight to the model, reasoning low81.67 points38.95 points85.28 points64.75 points30.26 points
GPT-6 LunaSent straight to the model, reasoning low84.2 points39.05 points84.95 points59.9 points28.51 points
Gemini 3.1 Flash-Lite (preview)Sent straight to the model, default setting85.48 points9.92 points89.46 points58.38 points48.38 points
Gemini 3.1 Flash-Lite (preview)Sent straight to the model, thinking medium83.43 points14.35 points87.67 points56.17 points49.05 points
Gemini 3.1 Flash-Lite (preview)Sent straight to the model, thinking high76.76 points30.89 points83.28 points56.61 points38.93 points
Gemini 3.5 Flash-LiteSent straight to the model, default setting73.48 points6.26 points82.22 points57.74 points66.03 points
GPT-5.6 LunaSent straight to the model, reasoning none81.3 points28.32 points82.65 points63.9 points25.43 points
Gemini 3.5 Flash-LiteSent straight to the model, thinking medium72.85 points6.23 points78.8 points48.05 points65.55 points
Claude Opus 4.6Sent straight to the model, default setting86.52 points13.49 points89.7 points64.19 points16.47 points
Claude Haiku 4.5Sent straight to the model, thinking on78.7 points27.39 points84.83 points62.36 points12.3 points
GPT-6 LunaSent straight to the model, reasoning none81.42 points15.48 points85.07 points59.84 points22.85 points
GPT-5 miniSent straight to the model, reasoning medium74.6 points38.96 points85.68 points45.35 points13.03 points
GPT-5 miniSent straight to the model, reasoning minimal69.82 points30.13 points82.3 points45.77 points6.15 points
Claude Haiku 4.5Sent straight to the model, thinking off77.21 points13.77 points78.74 points49.39 points6.72 points
GPT-5.4 NanoSent straight to the model, default setting60.16 points17.57 points78.05 points54.56 points6.41 points
Technical details for ParseBench

Overall: The average of five scores, each from 0 to 100: rebuilding tables, reading the numbers in charts, keeping the text complete and correct, keeping formatting that carries meaning such as headings, and placing each part of the page in the right spot.

  • Claude Opus 5.5, sent straight to the model, effort high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: effort high.
  • Claude Fable 5.1, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • Claude Opus 5.5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • GPT-5.6 Sol, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
  • Gemini 3 Flash (preview), sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3 Flash (preview), sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • Claude Opus 5.5, sent straight to the model, effort low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: effort low.
  • Gemini 3.5 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3.8 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3.7 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3 Flash (preview), sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
  • Claude Fable 5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • Gemini 3.7 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • Gemini 3.8 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • GPT-6 Astra, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
  • Gemini 3.8 Flash, sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
  • Gemini 3.6 Flash, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3.5 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • GPT-6 Sol, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
  • Gemini 3.7 Flash, sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
  • Gemini 3.1 Pro (preview), sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • GPT-5.6 Terra, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • GPT-5.6 Luna, sent straight to the model, reasoning max: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning max.
  • GPT-5.6 Luna, sent straight to the model, reasoning xhigh: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning xhigh.
  • GPT-6 Sol, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • GPT-5.6 Terra, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
  • GPT-5.5, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • GPT-5.6 Luna, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
  • Gemini 3.6 Flash, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • GPT-6 Sol, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • GPT-6 Luna, sent straight to the model, reasoning max: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning max.
  • Gemini 3.6 Flash, sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
  • GPT-5.5, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • GPT-6 Luna, sent straight to the model, reasoning xhigh: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning xhigh.
  • GPT-6 Luna, sent straight to the model, reasoning high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning high.
  • GPT-5.6 Terra, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • Claude Opus 4.8, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • GPT-5.6 Luna, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • Claude Opus 4.7, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • Gemini 3.5 Flash, sent straight to the model, thinking minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking minimal.
  • GPT-6 Luna, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • GPT-5.4, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • Claude Sonnet 5, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • GPT-5.6 Sol, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • Gemini 3.1 Pro (preview), sent straight to the model, thinking low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking low.
  • Gemini 3.5 Flash-Lite, sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • GPT-5.6 Luna, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
  • GPT-6 Luna, sent straight to the model, reasoning low: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning low.
  • Gemini 3.1 Flash-Lite (preview), sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • Gemini 3.1 Flash-Lite (preview), sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • Gemini 3.1 Flash-Lite (preview), sent straight to the model, thinking high: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking high.
  • Gemini 3.5 Flash-Lite, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • GPT-5.6 Luna, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • Gemini 3.5 Flash-Lite, sent straight to the model, thinking medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking medium.
  • Claude Opus 4.6, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.
  • Claude Haiku 4.5, sent straight to the model, thinking on: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking on.
  • GPT-6 Luna, sent straight to the model, reasoning none: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning none.
  • GPT-5 mini, sent straight to the model, reasoning medium: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning medium.
  • GPT-5 mini, sent straight to the model, reasoning minimal: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: reasoning minimal.
  • Claude Haiku 4.5, sent straight to the model, thinking off: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: thinking off.
  • GPT-5.4 Nano, sent straight to the model, default setting: ParseBench's own pipeline, run by LlamaIndex. Tools: None. Input: One page per request. Setting: default setting.

Source version: ParseBench leaderboard.csv at commit afb36bd, downloaded 2026-09-29

Results posted 2026-04-21 to 2026-09-25

Run dates not published by the tester; the dates are when results were posted.

License: Apache-2.0; LlamaIndex offered it for public use by email on 2026-09-27. Checked 2026-09-29.

IDP Core Bench

best 81.8%NanonetsNewest result 2026-03-23reliability not measuredcaveat

A test from Nanonets of how well AI reads everyday business documents: pulling out fields such as invoice numbers, dates and totals, reading printed and handwritten text, reading tables, and answering questions about a document.

Results

All Sent straight to the model, default settings · Run date not published; posted 2026-03-23

Each model's best setting in this test

  • Gemini 3 Pro (preview)81.8%
  • Claude Sonnet 4.681.2%
  • Claude Opus 4.681.1%
  • Gemini 3 Flash (preview)80.5%
  • GPT-5.277.4%
  • GPT-4.174.7%
  • GPT-5 mini73.3%
  • Claude Haiku 4.572.9%
Show all 9 results (1 not shown above) from IDP Core Bench
  • Gemini 3 Pro (preview)81.8%
  • Claude Sonnet 4.681.2%
  • Claude Opus 4.681.1%
  • Gemini 3 Flash (preview)80.5%
  • GPT-5.277.4%
  • GPT-4.174.7%
  • GPT-5 mini73.3%
  • Claude Haiku 4.572.9%
  • GPT-5 nano65.8%

Watch out

Nanonets does not publish when it ran each model. The date shown is when Nanonets last changed that model's result file, 2026-03-23 for every model here; the leaderboard page says "As of April 2026". Nanonets runs each model once per document with the provider's default settings and caps each answer at 8,192 tokens, which can cut off models that think at length. A request that failed, such as an image over a provider's size limit, counts as zero: 6 items for each Claude model shown and 96 for Gemini 3 Pro. Two models on the leaderboard, Gemini 3.1 Pro and GPT-5.4, are not shown: the page's scores for them differ from Nanonets' own result files, so they are held until Nanonets confirms which figures are right. The overall score averages four different tasks, and only one of them is pulling fields out of documents. Nanonets gives different sizes for the test on different pages (about 2,000 documents, 6,406 items, 5,376 items), so no single count is shown.

Reliability not measured by this source.

Nanonets built this benchmark and sells document-processing AI, including Nanonets OCR, which it tests alongside the general models. Only the general models are shown here.

See Nanonets's results
Not tested here: 22 models

Claude Fable 5, Claude Fable 5.1, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 5, Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.4, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: IDP Core Bench

What's in the test

Invoices, receipts, forms, handwritten pages, charts and scanned text, each sent to the model as an image. 5,376 of the items in Nanonets' result files are PNG images; the format of the rest is not published. The overall score is the average of four tasks: pulling out fields, reading text, reading tables, and answering questions about a document.

  • Invoices
  • Receipts
  • Forms
  • Handwritten documents
  • Charts

Where results change

By task

The four tasks the overall score averages. Each is scored its own way: fields, text and answers by how close they are to the correct text, tables by how well their rows, columns and cells match.

Model and setupPulling out fieldsReading textReading tablesAnswering questions about a document
Gemini 3 Pro (preview)Sent straight to the model, default settings85.7%81.8%95.8%64.1%
Claude Sonnet 4.6Sent straight to the model, default settings89.5%73.7%96.3%65.2%
Claude Opus 4.6Sent straight to the model, default settings89.8%74.0%96.0%64.4%
Gemini 3 Flash (preview)Sent straight to the model, default settings91.1%81.7%85.6%63.5%
GPT-5.2Sent straight to the model, default settings87.5%72.8%86.0%63.5%
GPT-4.1Sent straight to the model, default settings87.1%75.6%73.1%63.0%
GPT-5 miniSent straight to the model, default settings85.7%73.0%69.5%65.0%
Claude Haiku 4.5Sent straight to the model, default settings85.6%65.0%81.7%59.2%
GPT-5 nanoSent straight to the model, default settings84.7%69.6%45.3%63.5%
Technical details for IDP Core Bench

Overall: The overall score averages a model's results across four tasks: pulling key fields from documents, reading printed and handwritten text, reading tables, and answering questions about a document. Each answer is scored by how closely it matches the correct one, so this is not the share answered right.

  • Gemini 3 Pro (preview), sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • Claude Sonnet 4.6, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • Claude Opus 4.6, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • Gemini 3 Flash (preview), sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • GPT-5.2, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • GPT-4.1, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • GPT-5 mini, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • Claude Haiku 4.5, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.
  • GPT-5 nano, sent straight to the model, default settings: Nanonets' test pipeline, run by Nanonets. Tools: None. Input: One document image per request, with the task's prompt. Provider's default settings; answers capped at 8,192 tokens.

Source version: IDP Leaderboard page and Nanonets result file for Gemini-3-Pro (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Sonnet 4.6 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Opus 4.6 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Gemini-3-Flash (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5.2 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-4.1 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5-Mini (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for Claude Haiku 4.5 (last changed 2026-03-23), downloaded 2026-09-29; IDP Leaderboard page and Nanonets result file for GPT-5-Nano (last changed 2026-03-23), downloaded 2026-09-29

Results posted 2026-03-23

Run dates not published by the tester; the dates are when results were posted.

License: MIT, as the IDP Leaderboard site declares ("license":"https://opensource.org/licenses/MIT"); Nanonets agreed by email on 2026-09-28. Checked 2026-09-29.

OmniDocBench

best 92.91 pointsOpenDataLabNewest result 2026-03-31reliability not measuredcaveat

OmniDocBench tests how well a model reads a whole PDF page and turns it into Markdown. It checks the text, the tables and the formulas against a human-checked answer for each page.

Results

All Sent straight to the model · Run date not published; posted 2026-03-31

  • Gemini 3 Pro (preview)92.91 points
  • Gemini 3 Flash (preview)92.62 points
  • GPT-5.286.59 points

Watch out

OmniDocBench does not publish when it ran each model; the date shown is the dated update note on its project page for these models' evaluations. This page shows only the Overall figure; OmniDocBench also publishes separate text, table and formula sub-scores that are not shown here. Data: OmniDocBench, Apache 2.0.

Reliability not measured by this source.

OmniDocBench is run by OpenDataLab.

See OpenDataLab's results
Not tested here: 28 models

Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.4, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: OmniDocBench

What's in the test

The benchmark is a set of PDF pages covering many document types, layouts and languages. Every page has detailed labels for text, tables, formulas and reading order that people checked by hand.

  • Academic literature
  • Slides converted from PDF
  • Black and white books and textbooks
  • Colorful textbooks with images
  • Exam papers
  • Handwritten notes
  • Magazines
  • Research and financial reports
  • Newspapers
Technical details for OmniDocBench

Overall: OmniDocBench averages three parts on a 0 to 100 scale: text accuracy, table structure and formula accuracy.

  • Gemini 3 Pro (preview), sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.
  • Gemini 3 Flash (preview), sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.
  • GPT-5.2, sent straight to the model: OmniDocBench end-to-end parsing. Tools: None. Input: A PDF page image; the model returns Markdown. None.

Source version: OmniDocBench README at commit 9b46e6da431551535616273c0c57f4510c88b8c5

Results posted 2026-03-31

Run dates not published by the tester; the dates are when results were posted.

License: Apache License 2.0. Checked 2026-09-28.

No figures on this page for: Claude Opus 4.5, Claude Opus 5, Claude Sonnet 4.5, Claude Sonnet 5.5, Gemini 2.5 Pro, GPT-5.1, GPT-5.1 Codex, GPT-5.2 Codex, GPT-5.4 Mini.

Also tested by

Their figures are not shown here because permission to reuse them is not yet confirmed.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use