Writing and fixing code
Fixing real bugs and writing working code inside existing software projects. Good means the change passes the project's own tests.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- Reliability measured by a linked source; figures not shown here
- 5 sources, 35 models with figures
- Sources last checked 2026-09-28
- Ring 4: no rerun evidence yet
- SWE = SWE-bench Verified (SWE-bench), LB = LiveBench, TB = Terminal-Bench, CC = CCBench, ESWE = SWE-bench Verified (Epoch AI)
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| SWE-bench Verified (SWE-bench) | Claude Opus 4.5 Simple coding agent, reasoning effort highroughly tied | 76.8% | Gemini 3 Flash (preview) Simple coding agent, reasoning effort high | 75.8% |
| LiveBench | Claude Opus 5.5 Sent straight to the model | 89.3% | Claude Fable 5.1 Sent straight to the model | 86.4% |
| SWE-bench Verified (Epoch AI) | Claude Opus 4.7 Coding agent, effort max | 83.5% | Gemini 3.5 Flash Coding agent, effort high | 79.3% |
| Model | SWE-bench Verified (SWE-bench) | LiveBench | SWE-bench Verified (Epoch AI) |
|---|---|---|---|
| Claude Opus 4.5 | 1 of 11Simple coding agent, reasoning effort high | 14 of 28Sent straight to the model | |
| Claude Opus 4.6 | 3 of 11Simple coding agent, reasoning effort default | 19 of 28Sent straight to the model | 3 of 8Coding agent, effort default |
| Claude Opus 4.7 | 8 of 28Sent straight to the model | 1 of 8Coding agent, effort max | |
| Claude Sonnet 4.6 | 15 of 28Sent straight to the model | 6 of 8Coding agent, effort default | |
| Gemini 3 Flash (preview) | 2 of 11Simple coding agent, reasoning effort high | 5 of 8Coding agent, effort default | |
| Gemini 3 Pro (preview) | 4 of 11Simple coding agent, reasoning effort default | 7 of 8Coding agent, effort default | |
| Gemini 3.5 Flash | 19 of 28Sent straight to the model | 2 of 8Coding agent, effort high | |
| GPT-5.1 | 9 of 11Simple coding agent, reasoning effort medium | 8 of 8Coding agent, effort high | |
| GPT-5.2 | 5 of 11Simple coding agent, reasoning effort high | 24 of 28Sent straight to the model | |
| GPT-5.2 Codex | 5 of 11Simple coding agent, reasoning effort default | 5 of 28Sent straight to the model | |
| GPT-5.4 | 22 of 28Sent straight to the model | 4 of 8Coding agent, effort high |
The tests here put 6 pairs of models in a different order.
- In SWE-bench Verified (SWE-bench), Claude Opus 4.5, simple coding agent, reasoning effort high, placed above GPT-5.2 Codex, simple coding agent, reasoning effort default; in LiveBench, GPT-5.2 Codex, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Claude Sonnet 4.6, coding agent, effort default.
- In SWE-bench Verified (SWE-bench), Gemini 3 Flash (preview), simple coding agent, reasoning effort high, placed above Claude Opus 4.6, simple coding agent, reasoning effort default; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Gemini 3 Flash (preview), coding agent, effort default.
Show all 6 pairs placed in a different order
- In SWE-bench Verified (SWE-bench), Claude Opus 4.5, simple coding agent, reasoning effort high, placed above GPT-5.2 Codex, simple coding agent, reasoning effort default; in LiveBench, GPT-5.2 Codex, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Claude Sonnet 4.6, coding agent, effort default.
- In SWE-bench Verified (SWE-bench), Gemini 3 Flash (preview), simple coding agent, reasoning effort high, placed above Claude Opus 4.6, simple coding agent, reasoning effort default; in SWE-bench Verified (Epoch AI), Claude Opus 4.6, coding agent, effort default, placed above Gemini 3 Flash (preview), coding agent, effort default.
- In SWE-bench Verified (SWE-bench), Claude Opus 4.6, simple coding agent, reasoning effort default, placed above GPT-5.2 Codex, simple coding agent, reasoning effort default; in LiveBench, GPT-5.2 Codex, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in SWE-bench Verified (Epoch AI), Gemini 3.5 Flash, coding agent, effort high, placed above Claude Sonnet 4.6, coding agent, effort default.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in SWE-bench Verified (Epoch AI), GPT-5.4, coding agent, effort high, placed above Claude Sonnet 4.6, coding agent, effort default.
Findings for this category aren't written yet.
SWE-bench Verified (SWE-bench)
best 76.8%SWE-benchNewest result 2026-02-26reliability not measuredcaveat
A test of whether AI can fix real reported bugs in open-source software, checked by running the project's own tests.
Results
Each model's best setting in this test
- Claude Opus 4.5 Simple coding agent, reasoning effort high, 2026-02-1776.8%
- Gemini 3 Flash (preview) Simple coding agent, reasoning effort high, 2026-02-1775.8%
- Claude Opus 4.6 Simple coding agent, reasoning effort default, 2026-02-1775.6%
- Gemini 3 Pro (preview) Simple coding agent, reasoning effort default, 2025-11-1874.2%
- GPT-5.2 Simple coding agent, reasoning effort high, 2026-02-1772.8%
- GPT-5.2 Codex Simple coding agent, reasoning effort default, 2026-02-1972.8%
- Claude Sonnet 4.5 Simple coding agent, reasoning effort high, 2026-02-1771.4%
- Claude Haiku 4.5 Simple coding agent, reasoning effort high, 2026-02-1766.6%
Show all 15 results (7 not shown above) from SWE-bench Verified (SWE-bench)
- Claude Opus 4.5 Simple coding agent, reasoning effort high, 2026-02-1776.8%
- Gemini 3 Flash (preview) Simple coding agent, reasoning effort high, 2026-02-1775.8%
- Claude Opus 4.6 Simple coding agent, reasoning effort default, 2026-02-1775.6%
- Claude Opus 4.5 Simple coding agent, reasoning effort medium, 2025-11-2474.4%
- Gemini 3 Pro (preview) Simple coding agent, reasoning effort default, 2025-11-1874.2%
- GPT-5.2 Simple coding agent, reasoning effort high, 2026-02-1772.8%
- GPT-5.2 Codex Simple coding agent, reasoning effort default, 2026-02-1972.8%
- Claude Sonnet 4.5 Simple coding agent, reasoning effort high, 2026-02-1771.4%
- Claude Sonnet 4.5 Simple coding agent, reasoning effort default, 2025-09-2970.6%
- Gemini 3 Pro (preview) Simple coding agent, reasoning effort high, 2026-02-2669.6%
- GPT-5.2 Simple coding agent, reasoning effort default, 2025-12-1169.0%
- Claude Haiku 4.5 Simple coding agent, reasoning effort high, 2026-02-1766.6%
- GPT-5.1 Simple coding agent, reasoning effort medium, 2025-11-2066.0%
- GPT-5.1 Codex Simple coding agent, reasoning effort medium, 2025-11-2466.0%
- GPT-5 mini Simple coding agent, reasoning effort default, 2026-02-1756.2%
Watch out
OpenAI said in February 2026 that it stopped reporting this test. It found that leading models from several companies could reproduce some of the test's answers from memory, which suggests they saw them during training, and that many of the hardest problems have flawed tests. OpenAI makes models tested here, so this is a vendor's view, but the concern applies across companies. The test also covers only Python projects.
Reliability not measured by this source.
See SWE-bench's resultsNot tested here: 24 models
Claude Fable 5, Claude Fable 5.1, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: SWE-bench Verified (SWE-bench)
What's in the test
500 real problems reported on GitHub, from 12 open-source Python projects, each screened by experienced developers to make sure it is fair and solvable. The model gets the problem description and the project's code; its fix counts only if the project's own tests then pass.
Technical details for SWE-bench Verified (SWE-bench)
Resolved rate: The share of 500 real GitHub issues the model fixed so that the project's own tests pass. Shown here: SWE-bench's bash-only runs dated within 12 months of the date checked. They use a minimal agent; its version and reasoning effort are shown on each row.
- Claude Opus 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- Gemini 3 Flash (preview), simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- Claude Opus 4.6, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
- Claude Opus 4.5, simple coding agent, reasoning effort medium: mini-SWE-agent 1.16.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
- Gemini 3 Pro (preview), simple coding agent, reasoning effort default: mini-SWE-agent 1.15.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
- GPT-5.2, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- GPT-5.2 Codex, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
- Claude Sonnet 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- Claude Sonnet 4.5, simple coding agent, reasoning effort default: mini-SWE-agent 1.13.3, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
- Gemini 3 Pro (preview), simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- GPT-5.2, simple coding agent, reasoning effort default: mini-SWE-agent 1.17.2, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
- Claude Haiku 4.5, simple coding agent, reasoning effort high: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: high.
- GPT-5.1, simple coding agent, reasoning effort medium: mini-SWE-agent 1.15.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
- GPT-5.1 Codex, simple coding agent, reasoning effort medium: mini-SWE-agent 1.16.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: medium.
- GPT-5 mini, simple coding agent, reasoning effort default: mini-SWE-agent 2.0.0, bash only, run by SWE-agent. Tools: A bash shell only. Input: GitHub issue text and the repository. Reasoning effort: default.
Source version: SWE-bench Verified
Run 2025-09-29 to 2026-02-26
License: CC BY-NC 4.0 (website data), shown for non-commercial use with attribution. Checked 2026-09-25.
LiveBench
best 89.3%LiveBenchNewest result 2026-06-25reliability not measuredcaveat
A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.
Results
All Sent straight to the model · Run date not published; posted 2026-06-25
Each model's best setting in this test
- Claude Opus 5.589.3%
- Claude Fable 5.186.4%
- Claude Fable 586.0%
- GPT-5.6 Sol83.9%
- GPT-5.2 Codex83.6%
- GPT-5.6 Luna82.9%
- GPT-5.582.2%
- Claude Opus 4.782.1%
Show all 29 results (21 not shown above) from LiveBench
- Claude Opus 5.589.3%
- Claude Opus 5.589.3%
- Claude Fable 5.186.4%
- Claude Fable 586.0%
- GPT-5.6 Sol83.9%
- GPT-5.2 Codex83.6%
- GPT-5.6 Luna82.9%
- GPT-5.582.2%
- Claude Opus 4.782.1%
- Claude Opus 4.881.8%
- GPT-6 Sol81.8%
- Claude Opus 581.5%
- Claude Sonnet 580.7%
- GPT-6 Astra80.4%
- Claude Opus 4.579.7%
- Claude Sonnet 4.679.3%
- GPT-6 Luna79.0%
- Gemini 3.7 Flash78.9%
- GPT-5.6 Terra78.3%
- Claude Opus 4.678.2%
- Gemini 3.5 Flash78.2%
- Gemini 3.6 Flash77.9%
- GPT-5.477.5%
- Gemini 3.1 Pro (preview)76.5%
- Gemini 3.5 Flash-Lite76.1%
- GPT-5.276.1%
- Gemini 3.8 Flash72.5%
- GPT-5.4 Mini71.6%
- GPT-5.4 Nano70.8%
Watch out
LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.
Reliability not measured by this source.
LiveBench is funded by Abacus.AI.
See LiveBench's resultsNot tested here: 7 models
Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1, GPT-5.1 Codex.
More about this test: LiveBench
What's in the test
Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.
Technical details for LiveBench
Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
- GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
- GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
Source version: LiveBench release 2026-06-25
Results posted 2026-06-25
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-27.
Sources for this summary:
SWE-bench Verified (Epoch AI)
best 83.5%Epoch AI484 problemsNewest result 2026-06-01reliability not measuredcaveat
Epoch AI's own run of SWE-bench Verified, a test of whether AI can fix real reported bugs in open-source Python projects, using Epoch's own coding agent. These figures are separate from the SWE-bench team's own leaderboard.
Results
- Claude Opus 4.7 Coding agent, effort max, 2026-04-2083.5%
- Gemini 3.5 Flash Coding agent, effort high, 2026-06-0179.3%
- Claude Opus 4.6 Coding agent, effort default, 2026-02-1878.7%
- GPT-5.4 Coding agent, effort high, 2026-03-0676.9%
- Gemini 3 Flash (preview) Coding agent, effort default, 2026-02-1875.4%
- Claude Sonnet 4.6 Coding agent, effort default, 2026-02-2175.2%
- Gemini 3 Pro (preview) Coding agent, effort default, 2026-02-1372.9%
- GPT-5.1 Coding agent, effort high, 2026-02-1868.0%
Watch out
Epoch labels this test flawed and has estimated that 5 to 10% of its problems have errors. Epoch made a major upgrade to its agent and setup in February 2026, and scores rose significantly after it. We show only Epoch's runs from after its February 2026 agent upgrade, so every row here uses the same setup. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability not measured by this source.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.
See Epoch AI's resultsNot tested here: 27 models
Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.5, Claude Sonnet 5, Gemini 3.1 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5 mini, GPT-5.1 Codex, GPT-5.2, GPT-5.2 Codex, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: SWE-bench Verified (Epoch AI)
What's in the test
Epoch uses 484 of the test's 500 problems, each from a GitHub issue in one of 12 open-source Python projects. It leaves out 16 that do not run reliably on its setup. The model gets the issue and the code, and its fix counts only if the project's tests then pass.
484 problems
Technical details for SWE-bench Verified (Epoch AI)
Resolved rate: The share of Epoch AI's 484 problems where the model's fix made the project's tests pass, in Epoch AI's own runs.
- Claude Opus 4.7, coding agent, effort max: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: max.
- Gemini 3.5 Flash, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.
- Claude Opus 4.6, coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
- GPT-5.4, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.
- Gemini 3 Flash (preview), coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
- Claude Sonnet 4.6, coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
- Gemini 3 Pro (preview), coding agent, effort default: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: default.
- GPT-5.1, coding agent, effort high: Epoch AI's coding agent (Inspect). Tools: A bash shell, a file editor and a patch tool. Input: GitHub issue text and the repository. Effort: high.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2026-02-13 to 2026-06-01
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano.
Also tested by
Their figures are not shown here because permission to reuse them is not yet confirmed.
Terminal-Bench
Terminal-Bench (Stanford and Laude Institute), checked 2026-09-24The share of Terminal-Bench command-line tasks an agent resolves, run as repeated trials per task and shown with a 95% confidence interval.
Last updated 2026-09-21
See their resultsMore about this test: Terminal-Bench
Reliability: Terminal-Bench runs repeated trials of each task and reports the uncertainty of each resolution rate as a 95% confidence interval.
Technical details for Terminal-Bench
Resolution rate: The share of Terminal-Bench command-line tasks an agent resolves, run as repeated trials per task and shown with a 95% confidence interval.
License: Results license not stated. Checked 2026-09-24.
CCBench
CodeCrafters, checked 2026-09-24The share of tasks where the agent's code change passes the project's own hand-written tests, on tasks built from real, non-public codebases.
CCBench reports that it draws its coding tasks from private codebases so the tasks are unlikely to have appeared in a model's training data.
Last updated 2026-02-12
See their resultsMore about this test: CCBench
Reliability not measured by this source.
Technical details for CCBench
Success: The share of tasks where the agent's code change passes the project's own hand-written tests, on tasks built from real, non-public codebases.
License: Results license not stated. Checked 2026-09-24.
What this doesn't tell you
- Results depend heavily on the agent setup around the model, not only the model. Compare models only within the same setup.
- SWE-bench Verified uses open-source Python projects. Your codebase may behave differently.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- SWE-bench Verified (SWE-bench): Reliability not measured by this source.
- LiveBench: Reliability not measured by this source.
- SWE-bench Verified (Epoch AI): Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use