Assurance

Writing and fixing code

Fixing real bugs and writing working code inside existing software projects. Good means the change passes the project's own tests.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Agent works through the task with tools, as the words after it sayDirect sent straight to the modelFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, earned repeat runs, some or all models: Results checked by repeat runs, for some or all models, not yet

Vendor claims have not been collected for any category yet.

SWE 76.8%SWE 76.8%LB 89.3%LB 89.3%TB ↗CC ↗ESWE 83.5%ESWE 83.5%
ⓘ More about this category
  • Reliability measured by a linked source; figures not shown here
  • 5 sources, 35 models with figures
  • Sources last checked 2026-09-28
  • Ring 4: no rerun evidence yet
  • SWE = SWE-bench Verified (SWE-bench), LB = LiveBench, TB = Terminal-Bench, CC = CCBench, ESWE = SWE-bench Verified (Epoch AI)

Who does well

TestBestBest scoreNextNext score
SWE-bench Verified (SWE-bench)Claude Opus 4.5 Simple coding agent, reasoning effort highroughly tied76.8%Gemini 3 Flash (preview) Simple coding agent, reasoning effort high75.8%
LiveBenchClaude Opus 5.5 Sent straight to the model89.3%Claude Fable 5.1 Sent straight to the model86.4%
SWE-bench Verified (Epoch AI)Claude Opus 4.7 Coding agent, effort max83.5%Gemini 3.5 Flash Coding agent, effort high79.3%
Place within each test. Places are never added up across tests.
ModelSWE-bench Verified (SWE-bench)LiveBenchSWE-bench Verified (Epoch AI)
Claude Opus 4.51 of 11Simple coding agent, reasoning effort high14 of 28Sent straight to the model
Claude Opus 4.63 of 11Simple coding agent, reasoning effort default19 of 28Sent straight to the model3 of 8Coding agent, effort default
Claude Opus 4.78 of 28Sent straight to the model1 of 8Coding agent, effort max
Claude Sonnet 4.615 of 28Sent straight to the model6 of 8Coding agent, effort default
Gemini 3 Flash (preview)2 of 11Simple coding agent, reasoning effort high5 of 8Coding agent, effort default
Gemini 3 Pro (preview)4 of 11Simple coding agent, reasoning effort default7 of 8Coding agent, effort default
Gemini 3.5 Flash19 of 28Sent straight to the model2 of 8Coding agent, effort high
GPT-5.19 of 11Simple coding agent, reasoning effort medium8 of 8Coding agent, effort high
GPT-5.25 of 11Simple coding agent, reasoning effort high24 of 28Sent straight to the model
GPT-5.2 Codex5 of 11Simple coding agent, reasoning effort default5 of 28Sent straight to the model
GPT-5.422 of 28Sent straight to the model4 of 8Coding agent, effort high

The tests here put 6 pairs of models in a different order.

Show all 6 pairs placed in a different order
  • In SWE-bench Verified (SWE-bench), Claude Opus 4.5, simple coding agent, reasoning effort high, placed above GPT-5.2 Codex, simple coding agent, reasoning effort default; in LiveBench, GPT-5.2 Codex, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Claude Sonnet 4.6, coding agent, effort default.
  • In SWE-bench Verified (SWE-bench), Gemini 3 Flash (preview), simple coding agent, reasoning effort high, placed above Claude Opus 4.6, simple coding agent, reasoning effort default; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Gemini 3 Flash (preview), coding agent, effort default.
  • In SWE-bench Verified (SWE-bench), Claude Opus 4.6, simple coding agent, reasoning effort default, placed above GPT-5.2 Codex, simple coding agent, reasoning effort default; in LiveBench, GPT-5.2 Codex, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in SWE-bench Verified (Epoch AI), Gemini 3.5 Flash, coding agent, effort high, placed above Claude Sonnet 4.6, coding agent, effort default.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in SWE-bench Verified (Epoch AI), GPT-5.4, coding agent, effort high, placed above Claude Sonnet 4.6, coding agent, effort default.

Findings for this category aren't written yet.

The tests behind this

SWE-bench Verified (SWE-bench)

best 76.8%SWE-benchNewest result 2026-02-26reliability not measuredcaveat

A test of whether AI can fix real reported bugs in open-source software, checked by running the project's own tests.

Results

Each model's best setting in this test

  • Claude Opus 4.5 Simple coding agent, reasoning effort high, 2026-02-1776.8%
  • Gemini 3 Flash (preview) Simple coding agent, reasoning effort high, 2026-02-1775.8%
  • Claude Opus 4.6 Simple coding agent, reasoning effort default, 2026-02-1775.6%
  • Gemini 3 Pro (preview) Simple coding agent, reasoning effort default, 2025-11-1874.2%
  • GPT-5.2 Simple coding agent, reasoning effort high, 2026-02-1772.8%
  • GPT-5.2 Codex Simple coding agent, reasoning effort default, 2026-02-1972.8%
  • Claude Sonnet 4.5 Simple coding agent, reasoning effort high, 2026-02-1771.4%
  • Claude Haiku 4.5 Simple coding agent, reasoning effort high, 2026-02-1766.6%
Show all 15 results (7 not shown above) from SWE-bench Verified (SWE-bench)
  • Claude Opus 4.5 Simple coding agent, reasoning effort high, 2026-02-1776.8%
  • Gemini 3 Flash (preview) Simple coding agent, reasoning effort high, 2026-02-1775.8%
  • Claude Opus 4.6 Simple coding agent, reasoning effort default, 2026-02-1775.6%
  • Claude Opus 4.5 Simple coding agent, reasoning effort medium, 2025-11-2474.4%
  • Gemini 3 Pro (preview) Simple coding agent, reasoning effort default, 2025-11-1874.2%
  • GPT-5.2 Simple coding agent, reasoning effort high, 2026-02-1772.8%
  • GPT-5.2 Codex Simple coding agent, reasoning effort default, 2026-02-1972.8%
  • Claude Sonnet 4.5 Simple coding agent, reasoning effort high, 2026-02-1771.4%
  • Claude Sonnet 4.5 Simple coding agent, reasoning effort default, 2025-09-2970.6%
  • Gemini 3 Pro (preview) Simple coding agent, reasoning effort high, 2026-02-2669.6%
  • GPT-5.2 Simple coding agent, reasoning effort default, 2025-12-1169.0%
  • Claude Haiku 4.5 Simple coding agent, reasoning effort high, 2026-02-1766.6%
  • GPT-5.1 Simple coding agent, reasoning effort medium, 2025-11-2066.0%
  • GPT-5.1 Codex Simple coding agent, reasoning effort medium, 2025-11-2466.0%
  • GPT-5 mini Simple coding agent, reasoning effort default, 2026-02-1756.2%

Watch out

OpenAI said in February 2026 that it stopped reporting this test. It found that leading models from several companies could reproduce some of the test's answers from memory, which suggests they saw them during training, and that many of the hardest problems have flawed tests. OpenAI makes models tested here, so this is a vendor's view, but the concern applies across companies. The test also covers only Python projects.

Reliability not measured by this source.

See SWE-bench's results
Not tested here: 24 models

Claude Fable 5, Claude Fable 5.1, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: SWE-bench Verified (SWE-bench)

What's in the test

500 real problems reported on GitHub, from 12 open-source Python projects, each screened by experienced developers to make sure it is fair and solvable. The model gets the problem description and the project's code; its fix counts only if the project's own tests then pass.

Technical details for SWE-bench Verified (SWE-bench)

Resolved rate: The share of 500 real GitHub issues the model fixed so that the project's own tests pass. Shown here: SWE-bench's bash-only runs dated within 12 months of the date checked. They use a minimal agent; its version and reasoning effort are shown on each row.

  • Claude Opus 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • Gemini 3 Flash (preview), simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • Claude Opus 4.6, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
  • Claude Opus 4.5, simple coding agent, reasoning effort medium: mini-SWE-agent 1.16.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
  • Gemini 3 Pro (preview), simple coding agent, reasoning effort default: mini-SWE-agent 1.15.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
  • GPT-5.2, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • GPT-5.2 Codex, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
  • Claude Sonnet 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • Claude Sonnet 4.5, simple coding agent, reasoning effort default: mini-SWE-agent 1.13.3, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
  • Gemini 3 Pro (preview), simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • GPT-5.2, simple coding agent, reasoning effort default: mini-SWE-agent 1.17.2, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
  • Claude Haiku 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
  • GPT-5.1, simple coding agent, reasoning effort medium: mini-SWE-agent 1.15.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
  • GPT-5.1 Codex, simple coding agent, reasoning effort medium: mini-SWE-agent 1.16.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
  • GPT-5 mini, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.

Source version: SWE-bench Verified

Run 2025-09-29 to 2026-02-26

License: CC BY-NC 4.0 (website data), shown for non-commercial use with attribution. Checked 2026-09-25.

LiveBench

best 89.3%LiveBenchNewest result 2026-06-25reliability not measuredcaveat

A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.

Results

All Sent straight to the model · Run date not published; posted 2026-06-25

Each model's best setting in this test

  • Claude Opus 5.589.3%
  • Claude Fable 5.186.4%
  • Claude Fable 586.0%
  • GPT-5.6 Sol83.9%
  • GPT-5.2 Codex83.6%
  • GPT-5.6 Luna82.9%
  • GPT-5.582.2%
  • Claude Opus 4.782.1%
Show all 29 results (21 not shown above) from LiveBench
  • Claude Opus 5.589.3%
  • Claude Opus 5.589.3%
  • Claude Fable 5.186.4%
  • Claude Fable 586.0%
  • GPT-5.6 Sol83.9%
  • GPT-5.2 Codex83.6%
  • GPT-5.6 Luna82.9%
  • GPT-5.582.2%
  • Claude Opus 4.782.1%
  • Claude Opus 4.881.8%
  • GPT-6 Sol81.8%
  • Claude Opus 581.5%
  • Claude Sonnet 580.7%
  • GPT-6 Astra80.4%
  • Claude Opus 4.579.7%
  • Claude Sonnet 4.679.3%
  • GPT-6 Luna79.0%
  • Gemini 3.7 Flash78.9%
  • GPT-5.6 Terra78.3%
  • Claude Opus 4.678.2%
  • Gemini 3.5 Flash78.2%
  • Gemini 3.6 Flash77.9%
  • GPT-5.477.5%
  • Gemini 3.1 Pro (preview)76.5%
  • Gemini 3.5 Flash-Lite76.1%
  • GPT-5.276.1%
  • Gemini 3.8 Flash72.5%
  • GPT-5.4 Mini71.6%
  • GPT-5.4 Nano70.8%

Watch out

LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.

Reliability not measured by this source.

LiveBench is funded by Abacus.AI.

See LiveBench's results
Not tested here: 7 models

Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1, GPT-5.1 Codex.

More about this test: LiveBench

What's in the test

Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.

Technical details for LiveBench

Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.

  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
  • GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
  • GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.

Source version: LiveBench release 2026-06-25

Results posted 2026-06-25

Run dates not published by the tester; the dates are when results were posted.

License: Apache License 2.0. Checked 2026-09-27.

SWE-bench Verified (Epoch AI)

best 83.5%Epoch AI484 problemsNewest result 2026-06-01reliability not measuredcaveat

Epoch AI's own run of SWE-bench Verified, a test of whether AI can fix real reported bugs in open-source Python projects, using Epoch's own coding agent. These figures are separate from the SWE-bench team's own leaderboard.

Results

  • Claude Opus 4.7 Coding agent, effort max, 2026-04-2083.5%
  • Gemini 3.5 Flash Coding agent, effort high, 2026-06-0179.3%
  • Claude Opus 4.6 Coding agent, effort default, 2026-02-1878.7%
  • GPT-5.4 Coding agent, effort high, 2026-03-0676.9%
  • Gemini 3 Flash (preview) Coding agent, effort default, 2026-02-1875.4%
  • Claude Sonnet 4.6 Coding agent, effort default, 2026-02-2175.2%
  • Gemini 3 Pro (preview) Coding agent, effort default, 2026-02-1372.9%
  • GPT-5.1 Coding agent, effort high, 2026-02-1868.0%

Watch out

Epoch labels this test flawed and has estimated that 5 to 10% of its problems have errors. Epoch made a major upgrade to its agent and setup in February 2026, and scores rose significantly after it. We show only Epoch's runs from after its February 2026 agent upgrade, so every row here uses the same setup. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability not measured by this source.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.

See Epoch AI's results
Not tested here: 27 models

Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.5, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5 mini, GPT-5.1 Codex, GPT-5.2, GPT-5.2 Codex, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: SWE-bench Verified (Epoch AI)

What's in the test

Epoch uses 484 of the test's 500 problems, each from a GitHub issue in one of 12 open-source Python projects. It leaves out 16 that do not run reliably on its setup. The model gets the issue and the code, and its fix counts only if the project's tests then pass.

484 problems

Technical details for SWE-bench Verified (Epoch AI)

Resolved rate: The share of Epoch AI's 484 problems where the model's fix made the project's tests pass, in Epoch AI's own runs.

  • Claude Opus 4.7, coding agent, effort max: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: max.
  • Gemini 3.5 Flash, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.
  • Claude Opus 4.6, coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
  • GPT-5.4, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.
  • Gemini 3 Flash (preview), coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
  • Claude Sonnet 4.6, coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
  • Gemini 3 Pro (preview), coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
  • GPT-5.1, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2026-02-13 to 2026-06-01

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano.

Also tested by

Their figures are not shown here because permission to reuse them is not yet confirmed.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use