Assurance

Analyzing data and spreadsheets

Answering questions from tables and spreadsheets: totals, trends, combining tables and cleaning up messy data. Good means the numbers are exactly right.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Direct sent straight to the modelAgent works through the task with tools, as the words after it sayFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, earned repeat runs, some or all models: Results checked by repeat runs, for some or all models, not yet

Vendor claims have not been collected for any category yet.

LB 83.0%LB 83.0%LSQL 38.0%LSQL 38.0%
ⓘ More about this category
  • Reliability not yet measured
  • 2 sources, 29 models with figures
  • Sources last checked 2026-09-30
  • Ring 4: no rerun evidence yet
  • LB = LiveBench, LSQL = LiveSQLBench

Who does well

TestBestBest scoreNextNext score
LiveBenchGPT-6 Astra Sent straight to the model83.0%GPT-5.5 Sent straight to the model81.6%
LiveSQLBenchClaude Opus 4.6 Agent (OpenHands) working through a terminal only38.0%Claude Sonnet 4.5 Agent (OpenHands) working through a terminal only35.2%
Place within each test. Places are never added up across tests.
ModelLiveBenchLiveSQLBench
Claude Opus 4.621 of 28Sent straight to the model1 of 2Agent (OpenHands) working through a terminal only

Findings for this category aren't written yet.

The tests behind this

LiveBench

best 83.0%LiveBenchNewest result 2026-06-25reliability not measuredcaveat

A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.

Results

All Sent straight to the model · Run date not published; posted 2026-06-25

Each model's best setting in this test

  • GPT-6 Astra83.0%
  • GPT-5.581.6%
  • GPT-6 Sol81.2%
  • Claude Fable 580.5%
  • Claude Opus 5.580.3%
  • Claude Fable 5.180.3%
  • GPT-5.6 Sol79.8%
  • GPT-5.479.3%
  • GPT-5.6 Terra79.3%
Show all 29 results (20 not shown above) from LiveBench
  • GPT-6 Astra83.0%
  • GPT-5.581.6%
  • GPT-6 Sol81.2%
  • Claude Fable 580.5%
  • Claude Opus 5.580.3%
  • Claude Fable 5.180.3%
  • GPT-5.6 Sol79.8%
  • Claude Opus 5.579.8%
  • GPT-5.479.3%
  • GPT-5.6 Terra79.3%
  • Gemini 3.1 Pro (preview)78.5%
  • Claude Opus 4.778.3%
  • GPT-5.2 Codex78.2%
  • GPT-5.278.2%
  • GPT-5.6 Luna78.0%
  • Claude Sonnet 4.678.0%
  • Claude Opus 574.5%
  • Claude Opus 4.574.4%
  • GPT-6 Luna73.4%
  • Claude Sonnet 571.7%
  • GPT-5.4 Mini70.8%
  • Claude Opus 4.669.9%
  • Gemini 3.7 Flash68.0%
  • GPT-5.4 Nano67.6%
  • Claude Opus 4.866.0%
  • Gemini 3.5 Flash64.9%
  • Gemini 3.6 Flash63.0%
  • Gemini 3.8 Flash54.0%
  • Gemini 3.5 Flash-Lite53.3%

Watch out

LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.

Reliability not measured by this source.

LiveBench is funded by Abacus.AI.

See LiveBench's results
Not tested here: 1 model

Claude Sonnet 4.5.

More about this test: LiveBench

What's in the test

Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.

Technical details for LiveBench

Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.

  • GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
  • GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
  • Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.

Source version: LiveBench release 2026-06-25

Results posted 2026-06-25

Run dates not published by the tester; the dates are when results were posted.

License: Apache License 2.0. Checked 2026-09-27.

LiveSQLBench

best 38.0%BIRD team (University of Hong Kong)600 tasksNewest result 2026-03-27reliability not measuredcaveat

LiveSQLBench, from the BIRD team, tests whether AI can turn a business question into working SQL on real-world style databases, and also carry out database changes such as creating, updating and deleting records.

Results

All Agent (OpenHands) working through a terminal only · Dated 2026-03-27; LiveSQLBench does not say whether this is when it ran or when it was added

  • Claude Opus 4.638.0%
  • Claude Sonnet 4.535.2%

Watch out

Only results the BIRD team ran itself are shown. Several Claude, GPT and Gemini results on its site are held back: for some the site does not say who ran them, and others were sent in by an outside party and rerun by the BIRD team, which reports the best of several runs. The date shown is the one LiveSQLBench lists for each result; it does not say whether that is the day the model was run or the day the result was added. Each model was run once, as far as the BIRD team publishes. LiveSQLBench's own headline view blends two releases, so its numbers differ from the Base-Full v1 figures shown here. No cost is published for the results shown. Data: BIRD team, LiveSQLBench, shown with the BIRD team's permission.

Reliability not measured by this source.

The BIRD team at the University of Hong Kong built LiveSQLBench with Google Cloud, which it credits as a co-creator. Google makes Gemini, one of the model families on the leaderboard. No Gemini result is shown here.

See BIRD team (University of Hong Kong)'s results
Not tested here: 27 models

Claude Fable 5, Claude Fable 5.1, Claude Opus 4.5, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.2, GPT-5.2 Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: LiveSQLBench

What's in the test

Base-Full v1 has 600 tasks over 22 databases with messy, realistic data, each with a knowledge base of business rules the model must use. A task passes when the query returns the right result or the database change passes the BIRD team's tests. The BIRD team publishes new releases over time; only Base-Full v1 is shown here, because releases are different tests.

600 tasks

  • Questions answered with SQL
  • Database changes
  • Business rules from a knowledge base
Technical details for LiveSQLBench

Success rate: The share of the 600 Base-Full v1 tasks where the model's SQL passed the BIRD team's test cases.

  • Claude Opus 4.6, agent (OpenHands) working through a terminal only: OpenHands, through LiveSQLBench-CLI, run by the BIRD team. Tools: A terminal with the database, in a sandbox. Input: The question and the database's knowledge base, explored through the terminal. Effort not stated.
  • Claude Sonnet 4.5, agent (OpenHands) working through a terminal only: OpenHands, through LiveSQLBench-CLI, run by the BIRD team. Tools: A terminal with the database, in a sandbox. Input: The question and the database's knowledge base, explored through the terminal. Effort not stated.

Source version: LiveSQLBench leaderboard data (page-a0bf63f2d70a35a4.js), downloaded 2026-09-30

Dated 2026-03-27

Run dates not published; LiveSQLBench does not say whether these dates are when it ran or when results were added.

License: Display permitted by the BIRD team in writing, 2026-09-28; task data CC BY 4.0. Checked 2026-09-30.

No figures on this page for: Claude Haiku 4.5, Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 mini, GPT-5 nano, GPT-5.1, GPT-5.1 Codex.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use