Findings about setups from every category. Cost appears only where the tester published it for the same run.
Updated every Friday. Sources last checked 2026-09-26.
Pulling data from documents
In ExtractBench (LlamaIndex)
On long documents in ExtractBench (LlamaIndex), the best agent setup scored 88.1% and the best single request scored 35.8%, a gap of 52.3 points.
This tester publishes cost for the whole test, not for this part of it.
See the evidenceIn ExtractBench (LlamaIndex)
On short documents in ExtractBench (LlamaIndex), asking the model to show where each value came from changed accuracy by at most 0.5 points (2 models tried both ways).
This tester publishes cost for the whole test, not for this part of it.
See the evidenceIn ExtractBench (LlamaIndex)
On long documents it was different: asking for sources changed accuracy by as much as 10.1 points, for Claude Opus 4.8, coding agent (Claude Code); and GPT-5.5, coding agent (Codex).
This tester publishes cost for the whole test, not for this part of it.
See the evidence
Data version 2026-09-30+832fb354fcbe · Terms of use