Analyzing data and spreadsheets
Answering questions from tables and spreadsheets: totals, trends, combining tables and cleaning up messy data. Good means the numbers are exactly right.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- Reliability not yet measured
- 2 sources, 29 models with figures
- Sources last checked 2026-09-30
- Ring 4: no rerun evidence yet
- LB = LiveBench, LSQL = LiveSQLBench
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| LiveBench | GPT-6 Astra Sent straight to the model | 83.0% | GPT-5.5 Sent straight to the model | 81.6% |
| LiveSQLBench | Claude Opus 4.6 Agent (OpenHands) working through a terminal only | 38.0% | Claude Sonnet 4.5 Agent (OpenHands) working through a terminal only | 35.2% |
| Model | LiveBench | LiveSQLBench |
|---|---|---|
| Claude Opus 4.6 | 21 of 28Sent straight to the model | 1 of 2Agent (OpenHands) working through a terminal only |
Findings for this category aren't written yet.
LiveBench
best 83.0%LiveBenchNewest result 2026-06-25reliability not measuredcaveat
A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.
Results
All Sent straight to the model · Run date not published; posted 2026-06-25
Each model's best setting in this test
- GPT-6 Astra83.0%
- GPT-5.581.6%
- GPT-6 Sol81.2%
- Claude Fable 580.5%
- Claude Opus 5.580.3%
- Claude Fable 5.180.3%
- GPT-5.6 Sol79.8%
- GPT-5.479.3%
- GPT-5.6 Terra79.3%
Show all 29 results (20 not shown above) from LiveBench
- GPT-6 Astra83.0%
- GPT-5.581.6%
- GPT-6 Sol81.2%
- Claude Fable 580.5%
- Claude Opus 5.580.3%
- Claude Fable 5.180.3%
- GPT-5.6 Sol79.8%
- Claude Opus 5.579.8%
- GPT-5.479.3%
- GPT-5.6 Terra79.3%
- Gemini 3.1 Pro (preview)78.5%
- Claude Opus 4.778.3%
- GPT-5.2 Codex78.2%
- GPT-5.278.2%
- GPT-5.6 Luna78.0%
- Claude Sonnet 4.678.0%
- Claude Opus 574.5%
- Claude Opus 4.574.4%
- GPT-6 Luna73.4%
- Claude Sonnet 571.7%
- GPT-5.4 Mini70.8%
- Claude Opus 4.669.9%
- Gemini 3.7 Flash68.0%
- GPT-5.4 Nano67.6%
- Claude Opus 4.866.0%
- Gemini 3.5 Flash64.9%
- Gemini 3.6 Flash63.0%
- Gemini 3.8 Flash54.0%
- Gemini 3.5 Flash-Lite53.3%
Watch out
LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.
Reliability not measured by this source.
LiveBench is funded by Abacus.AI.
See LiveBench's resultsNot tested here: 1 model
Claude Sonnet 4.5.
More about this test: LiveBench
What's in the test
Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.
Technical details for LiveBench
Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.
- GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
- GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
- Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
Source version: LiveBench release 2026-06-25
Results posted 2026-06-25
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-27.
Sources for this summary:
LiveSQLBench
best 38.0%BIRD team (University of Hong Kong)600 tasksNewest result 2026-03-27reliability not measuredcaveat
LiveSQLBench, from the BIRD team, tests whether AI can turn a business question into working SQL on real-world style databases, and also carry out database changes such as creating, updating and deleting records.
Results
All Agent (OpenHands) working through a terminal only · Dated 2026-03-27; LiveSQLBench does not say whether this is when it ran or when it was added
- Claude Opus 4.638.0%
- Claude Sonnet 4.535.2%
Watch out
Only results the BIRD team ran itself are shown. Several Claude, GPT and Gemini results on its site are held back: for some the site does not say who ran them, and others were sent in by an outside party and rerun by the BIRD team, which reports the best of several runs. The date shown is the one LiveSQLBench lists for each result; it does not say whether that is the day the model was run or the day the result was added. Each model was run once, as far as the BIRD team publishes. LiveSQLBench's own headline view blends two releases, so its numbers differ from the Base-Full v1 figures shown here. No cost is published for the results shown. Data: BIRD team, LiveSQLBench, shown with the BIRD team's permission.
Reliability not measured by this source.
The BIRD team at the University of Hong Kong built LiveSQLBench with Google Cloud, which it credits as a co-creator. Google makes Gemini, one of the model families on the leaderboard. No Gemini result is shown here.
See BIRD team (University of Hong Kong)'s resultsNot tested here: 27 models
Claude Fable 5, Claude Fable 5.1, Claude Opus 4.5, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.2, GPT-5.2 Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: LiveSQLBench
What's in the test
Base-Full v1 has 600 tasks over 22 databases with messy, realistic data, each with a knowledge base of business rules the model must use. A task passes when the query returns the right result or the database change passes the BIRD team's tests. The BIRD team publishes new releases over time; only Base-Full v1 is shown here, because releases are different tests.
600 tasks
- Questions answered with SQL
- Database changes
- Business rules from a knowledge base
Technical details for LiveSQLBench
Success rate: The share of the 600 Base-Full v1 tasks where the model's SQL passed the BIRD team's test cases.
- Claude Opus 4.6, agent (OpenHands) working through a terminal only: OpenHands, through LiveSQLBench-CLI, run by the BIRD team. Tools: A terminal with the database, in a sandbox. Input: The question and the database's knowledge base, explored through the terminal. Effort not stated.
- Claude Sonnet 4.5, agent (OpenHands) working through a terminal only: OpenHands, through LiveSQLBench-CLI, run by the BIRD team. Tools: A terminal with the database, in a sandbox. Input: The question and the database's knowledge base, explored through the terminal. Effort not stated.
Source version: LiveSQLBench leaderboard data (page-a0bf63f2d70a35a4.js), downloaded 2026-09-30
Dated 2026-03-27
Run dates not published; LiveSQLBench does not say whether these dates are when it ran or when results were added.
License: Display permitted by the BIRD team in writing, 2026-09-28; task data CC BY 4.0. Checked 2026-09-30.
No figures on this page for: Claude Haiku 4.5, Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.1, GPT-5.1 Codex.
What this doesn't tell you
- None of these tests used your data.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- LiveBench: Reliability not measured by this source.
- LiveSQLBench: Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use