Assurance

Scientific tasks

Scientific work, such as answering expert-level questions in biology, chemistry or physics, reading research papers, or analyzing experimental data. A good result is correct and matches what an expert would conclude.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Direct sent straight to the modelFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, not yet repeat runs, some or all models: Results checked by repeat runs, for some or all models, not yet

Vendor claims have not been collected for any category yet.

GPQA 95.8%GPQA 95.8%
ⓘ More about this category
  • A source here repeats runs for some or all models; its method is under the source. Per-model figures not shown.
  • 1 source, 29 models with figures
  • Sources last checked 2026-09-28
  • Ring 4: needs rings 1 to 3 first
  • GPQA = GPQA Diamond

Who does well

TestBestBest scoreNextNext score
GPQA DiamondGPT-6 Astra Sent straight to the model, effort maxroughly tied, 3 setupsAlso roughly tied: Gemini 3.7 Flash Sent straight to the model, effort high 94.8%95.8%Gemini 3.8 Flash Sent straight to the model, effort high95.4%

Findings for this category aren't written yet.

The tests behind this

GPQA Diamond

best 95.8%Epoch AI198 questionsNewest result 2026-09-02repeats runs, method under the sourcecaveat

Hard multiple-choice science questions in biology, chemistry and physics, written by experts who hold or are working toward PhDs. Epoch AI runs the test on each model itself.

Results

Each model's best setting in this test

  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-3095.8%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0295.4%
  • Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1494.8%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0694.4%
  • Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0294.1%
  • Claude Opus 5 Sent straight to the model, effort max, 2026-07-2493.9%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0993.5%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0993.3%
Show all 78 results (70 not shown above) from GPQA Diamond
  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-3095.8%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0295.4%
  • Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1494.8%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0694.4%
  • Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0294.1%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-2094.1%
  • Claude Opus 5 Sent straight to the model, effort max, 2026-07-2493.9%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0993.5%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0993.3%
  • GPT-5.4 Sent straight to the model, effort xhigh, 2026-03-0693.3%
  • Claude Opus 5 Sent straight to the model, effort default, 2026-08-0692.9%
  • Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2292.8%
  • Gemini 3 Pro (preview) Sent straight to the model, effort default, 2025-11-1992.6%
  • GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0991.6%
  • GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1391.4%
  • Claude Opus 4.8 Sent straight to the model, effort max, 2026-06-0791.0%
  • GPT-5.5 Sent straight to the model, effort low, 2026-05-0590.7%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 32K, 2026-02-0690.5%
  • Claude Sonnet 5 Sent straight to the model, effort xhigh, 2026-07-0190.5%
  • Claude Opus 4.7 Sent straight to the model, effort xhigh, 2026-04-1790.2%
  • GPT-5.4 Sent straight to the model, effort high, 2026-07-1589.9%
  • GPT-5.6 Sol Sent straight to the model, effort low, 2026-08-0789.9%
  • Gemini 3 Flash (preview) Sent straight to the model, effort high, 2026-08-0689.4%
  • Gemini 3.5 Flash Sent straight to the model, effort low, 2026-08-0688.9%
  • GPT-5.4 Sent straight to the model, effort medium, 2026-07-1588.9%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 64K, 2026-02-0688.8%
  • Claude Opus 4.6 Sent straight to the model, effort max, 2026-08-0688.4%
  • Claude Opus 4.8 Sent straight to the model, effort low, 2026-08-0688.4%
  • GPT-5.2 Sent straight to the model, effort high, 2025-12-1188.2%
  • Claude Opus 5 Sent straight to the model, effort low, 2026-08-0687.9%
  • GPT-5.2 Sent straight to the model, effort medium, 2025-12-1187.9%
  • GPT-5.1 Sent straight to the model, effort high, 2025-11-1387.6%
  • Claude Sonnet 4.6 Sent straight to the model, thinking budget 32K, 2026-02-2087.4%
  • GPT-5.6 Terra Sent straight to the model, effort low, 2026-08-0787.4%
  • GPT-5.4 Mini Sent straight to the model, effort xhigh, 2026-08-0786.9%
  • Claude Opus 4.7 Sent straight to the model, effort max, 2026-08-0686.4%
  • Gemini 3.5 Flash Sent straight to the model, effort minimal, 2026-07-1586.4%
  • Gemini 3.6 Flash Sent straight to the model, effort low, 2026-08-0686.4%
  • Claude Opus 4.5 Sent straight to the model, thinking budget 32K, 2025-11-2486.0%
  • Claude Fable 5 Sent straight to the model, effort max, 2026-08-0685.9%
  • Gemini 3.6 Flash Sent straight to the model, effort minimal, 2026-08-0785.9%
  • Claude Opus 4.5 Sent straight to the model, thinking budget 16K, 2025-11-2585.5%
  • Claude Opus 4.8 Sent straight to the model, effort none, 2026-08-0685.4%
  • GPT-5.1 Sent straight to the model, effort medium, 2025-11-1785.0%
  • GPT-5.4 Sent straight to the model, effort low, 2026-07-1584.8%
  • GPT-5.4 Mini Sent straight to the model, effort high, 2026-04-1583.6%
  • Claude Fable 5 Sent straight to the model, effort high, 2026-08-0683.3%
  • Claude Sonnet 4.6 Sent straight to the model, effort high, 2026-07-1383.3%
  • Claude Sonnet 4.6 Sent straight to the model, effort medium, 2026-07-1383.3%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort high, 2026-08-0683.3%
  • Gemini 3 Flash (preview) Sent straight to the model, effort default, 2025-12-1783.2%
  • GPT-5.6 Sol Sent straight to the model, effort none, 2026-08-0782.8%
  • GPT-5.2 Sent straight to the model, effort low, 2025-12-1182.7%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 59K, 2025-10-2882.3%
  • GPT-5.6 Luna Sent straight to the model, effort low, 2026-08-0782.3%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2181.7%
  • Claude Opus 4.5 Sent straight to the model, effort default, 2025-11-2480.7%
  • Claude Sonnet 5 Sent straight to the model, effort max, 2026-08-0680.3%
  • Claude Fable 5 Sent straight to the model, effort low, 2026-08-0678.8%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 16K, 2025-10-2878.8%
  • Claude Sonnet 4.6 Sent straight to the model, effort max, 2026-08-0678.8%
  • GPT-5.4 Nano Sent straight to the model, effort high, 2026-04-1478.5%
  • GPT-5.5 Sent straight to the model, effort none, 2026-08-0777.3%
  • GPT-5.6 Terra Sent straight to the model, effort none, 2026-08-0777.3%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort low, 2026-08-0675.8%
  • GPT-5 mini Sent straight to the model, effort high, 2025-10-3075.0%
  • GPT-5.4 Sent straight to the model, effort none, 2026-07-1574.7%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort minimal, 2026-08-0674.2%
  • Claude Sonnet 4.5 Sent straight to the model, effort default, 2025-09-2973.7%
  • GPT-5.2 Sent straight to the model, effort none, 2026-07-1373.2%
  • GPT-5.4 Nano Sent straight to the model, effort low, 2026-08-0772.2%
  • GPT-5 mini Sent straight to the model, effort minimal, 2026-08-0771.7%
  • Claude Haiku 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2271.2%
  • GPT-5.1 Sent straight to the model, effort none, 2026-08-0766.7%
  • GPT-5.4 Mini Sent straight to the model, effort none, 2026-08-0764.1%
  • GPT-5.6 Luna Sent straight to the model, effort none, 2026-08-0763.6%
  • Claude Haiku 4.5 Sent straight to the model, effort default, 2025-10-1660.5%
  • GPT-5.4 Nano Sent straight to the model, effort none, 2026-08-0755.6%

Watch out

An answer not given in the required format gets no credit, so a model that formats answers badly can score below the 25% guessing rate. OpenAI found that PhD-level experts scored 69.7% on this set. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability: Epoch AI says it runs most models 16 times on this test and reports the average score. How many runs each model got is not in its download, and a few scores do not fit 16 runs.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.

See Epoch AI's results
More about this test: GPQA Diamond

What's in the test

The Diamond part of GPQA: 198 questions that both expert reviewers answered correctly but most non-experts got wrong. Each question has four options, so guessing scores about 25%.

198 questions

Older results, more than a year old (1) from GPQA Diamond
  • GPT-5 mini Sent straight to the model, effort medium, 2025-08-0771.7%
Technical details for GPQA Diamond

Accuracy: The share of the 198 questions the model answered correctly, in Epoch AI's own runs.

  • GPT-6 Astra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.8 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.7 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.6 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Opus 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Sol, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Terra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.4, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Opus 5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Gemini 3.5 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.6 Luna, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.2, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Opus 4.8, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Sonnet 5, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Opus 4.7, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • GPT-5.4, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.6 Sol, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Gemini 3 Flash (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.5 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Claude Opus 4.6, sent straight to the model, thinking budget 64K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 64K tokens.
  • Claude Opus 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Opus 4.8, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.2, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.2, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • GPT-5.1, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Sonnet 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5.6 Terra, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4 Mini, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Opus 4.7, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.5 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Gemini 3.6 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Fable 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.6 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Claude Opus 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
  • Claude Opus 4.8, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.1, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • GPT-5.4, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4 Mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Fable 5, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Sonnet 4.6, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Sonnet 4.6, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3 Flash (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.6 Sol, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.2, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 59K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 59K tokens.
  • GPT-5.6 Luna, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Opus 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Sonnet 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Fable 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
  • Claude Sonnet 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.4 Nano, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.5, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.6 Terra, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5 mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Claude Sonnet 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.2, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.4 Nano, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5 mini, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Claude Haiku 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5.1, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.4 Mini, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.6 Luna, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Haiku 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.4 Nano, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2025-09-29 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

No figures on this page for: Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano, GPT-5.1 Codex, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use