Mathematics
Solving mathematical problems, such as working out quantities in a word problem, checking a calculation, or finding the steps of a proof. A good result reaches the correct answer, not just a plausible one.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- A source here repeats runs for some or all models; its method is under the source. Per-model figures not shown.
- 5 sources, 34 models with figures
- Sources last checked 2026-09-30
- Ring 4: the tester says it repeats its runs; its method is under the source.
- LB = LiveBench, FM = FrontierMath (tiers 1 to 3), FM4 = FrontierMath (tier 4), AIME = OTIS Mock AIME, AXM = MathArena ArXivMath (June 2026)
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| LiveBench | Claude Opus 5.5 Sent straight to the modelroughly tied, 5 setupsAlso roughly tied: GPT-6 Astra Sent straight to the model 96.8%; GPT-6 Sol Sent straight to the model 96.4%; GPT-5.6 Sol Sent straight to the model 96.2% | 97.1% | Claude Fable 5.1 Sent straight to the model | 97.0% |
| FrontierMath (tiers 1 to 3) | GPT-6 Astra Model with a Python tool, effort max | 93.7% | Claude Fable 5.1 Model with a Python tool, effort max | 90.2% |
| FrontierMath (tier 4) | GPT-6 Astra Model with a Python tool, effort high | 97.6% | Claude Fable 5 Model with a Python tool, effort max | 90.2% |
| OTIS Mock AIME | Claude Fable 5 Sent straight to the model, effort highroughly tied, 5 setupsAlso roughly tied: GPT-5.6 Sol Sent straight to the model, effort max 100.0%; GPT-6 Astra Sent straight to the model, effort max 100.0%; GPT-5.6 Terra Sent straight to the model, effort max 99.7% | 100.0% | Claude Fable 5.1 Sent straight to the model, effort max | 100.0% |
| MathArena ArXivMath (June 2026) | Claude Fable 5 Sent straight to the model, effort max | 85.4% | GPT-5.5 Sent straight to the model, effort xhigh | 83.6% |
| Model | LiveBench | FrontierMath (tiers 1 to 3) | FrontierMath (tier 4) | OTIS Mock AIME | MathArena ArXivMath (June 2026) |
|---|---|---|---|---|---|
| Claude Fable 5 | 6 of 28Sent straight to the model | 4 of 26Model with a Python tool, effort max | 2 of 26Model with a Python tool, effort max | 1 of 30Sent straight to the model, effort high | 1 of 4Sent straight to the model, effort max |
| Claude Fable 5.1 | 2 of 28Sent straight to the model | 2 of 26Model with a Python tool, effort max | 3 of 26Model with a Python tool, effort max | 1 of 30Sent straight to the model, effort max | |
| Claude Opus 4.5 | 19 of 28Sent straight to the model | 24 of 26Model with a Python tool, thinking budget 32K | 24 of 26Model with a Python tool, thinking budget 32K | 25 of 30Sent straight to the model, thinking budget 32K | |
| Claude Opus 4.6 | 20 of 28Sent straight to the model | 15 of 26Model with a Python tool, effort max | 15 of 26Model with a Python tool, effort max | 18 of 30Sent straight to the model, thinking budget 64K | |
| Claude Opus 4.7 | 15 of 28Sent straight to the model | 12 of 26Model with a Python tool, effort max | 12 of 26Model with a Python tool, effort max | 10 of 30Sent straight to the model, effort xhigh | |
| Claude Opus 4.8 | 10 of 28Sent straight to the model | 9 of 26Model with a Python tool, effort max | 9 of 26Model with a Python tool, effort max | 8 of 30Sent straight to the model, effort max | |
| Claude Opus 5 | 8 of 28Sent straight to the model | 6 of 26Model with a Python tool, effort max | 5 of 26Model with a Python tool, effort max | 6 of 30Sent straight to the model, effort max | |
| Claude Sonnet 4.5 | 26 of 26Model with a Python tool, thinking budget 32K | 25 of 26Model with a Python tool, thinking budget 32K | 28 of 30Sent straight to the model, thinking budget 32K | ||
| Claude Sonnet 4.6 | 25 of 28Sent straight to the model | 26 of 30Sent straight to the model, thinking budget 32K | |||
| Claude Sonnet 5 | 14 of 28Sent straight to the model | 16 of 26Model with a Python tool, effort max | 14 of 26Model with a Python tool, effort max | 17 of 30Sent straight to the model, effort xhigh | |
| Gemini 3 Flash (preview) | 20 of 26Model with a Python tool, effort default | 20 of 26Model with a Python tool, effort default | 15 of 30Sent straight to the model, effort high | ||
| Gemini 3.1 Pro (preview) | 17 of 28Sent straight to the model | 18 of 26Model with a Python tool, effort default | 15 of 26Model with a Python tool, effort default | 14 of 30Sent straight to the model, effort default | 3 of 4Sent straight to the model, effort not stated |
| Gemini 3.5 Flash | 23 of 28Sent straight to the model | 17 of 26Model with a Python tool, effort high | 15 of 26Model with a Python tool, effort high | 15 of 30Sent straight to the model, effort high | 4 of 4Sent straight to the model, effort not stated |
| Gemini 3.5 Flash-Lite | 28 of 28Sent straight to the model | 25 of 26Model with a Python tool, effort high | 26 of 26Model with a Python tool, effort high | 29 of 30Sent straight to the model, effort high | |
| Gemini 3.6 Flash | 26 of 28Sent straight to the model | 19 of 26Model with a Python tool, effort high | 18 of 26Model with a Python tool, effort high | 19 of 30Sent straight to the model, effort high | |
| Gemini 3.7 Flash | 12 of 28Sent straight to the model | 11 of 26Model with a Python tool, effort high | 11 of 26Model with a Python tool, effort high | 12 of 30Sent straight to the model, effort high | |
| Gemini 3.8 Flash | 16 of 28Sent straight to the model | 13 of 26Model with a Python tool, effort high | 18 of 26Model with a Python tool, effort high | 6 of 30Sent straight to the model, effort high | |
| GPT-5 mini | 22 of 26Model with a Python tool, effort high | 21 of 26Model with a Python tool, effort high | 24 of 30Sent straight to the model, effort high | ||
| GPT-5.2 | 13 of 28Sent straight to the model | 14 of 26Model with a Python tool, effort xhigh | 13 of 26Model with a Python tool, effort xhigh | 13 of 30Sent straight to the model, effort high | |
| GPT-5.4 | 11 of 28Sent straight to the model | 10 of 26Model with a Python tool, effort xhigh | 10 of 26Model with a Python tool, effort xhigh | 11 of 30Sent straight to the model, effort high | |
| GPT-5.4 Mini | 27 of 28Sent straight to the model | 20 of 26Model with a Python tool, effort xhigh | 23 of 26Model with a Python tool, effort xhigh | 21 of 30Sent straight to the model, effort xhigh | |
| GPT-5.4 Nano | 18 of 28Sent straight to the model | 23 of 26Model with a Python tool, effort high | 21 of 26Model with a Python tool, effort high | 23 of 30Sent straight to the model, effort high | |
| GPT-5.5 | 7 of 28Sent straight to the model | 7 of 26Model with a Python tool, effort xhigh | 6 of 26Model with a Python tool, effort xhigh | 27 of 30Sent straight to the model, effort low | 2 of 4Sent straight to the model, effort xhigh |
| GPT-5.6 Luna | 24 of 28Sent straight to the model | 8 of 26Model with a Python tool, effort max | 8 of 26Model with a Python tool, effort max | 8 of 30Sent straight to the model, effort max | |
| GPT-5.6 Sol | 5 of 28Sent straight to the model | 3 of 26Model with a Python tool, effort max | 4 of 26Model with a Python tool, effort max | 1 of 30Sent straight to the model, effort max | |
| GPT-5.6 Terra | 9 of 28Sent straight to the model | 5 of 26Model with a Python tool, effort max | 7 of 26Model with a Python tool, effort max | 5 of 30Sent straight to the model, effort max | |
| GPT-6 Astra | 3 of 28Sent straight to the model | 1 of 26Model with a Python tool, effort max | 1 of 26Model with a Python tool, effort high | 1 of 30Sent straight to the model, effort max |
The tests here put 200 pairs of models in a different order.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In FrontierMath (tiers 1 to 3), Claude Fable 5.1, model with a Python tool, effort max, placed above Claude Fable 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above GPT-5.6 Sol, model with a Python tool, effort max.
Show all 200 pairs placed in a different order
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In FrontierMath (tiers 1 to 3), Claude Fable 5.1, model with a Python tool, effort max, placed above Claude Fable 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above GPT-5.6 Sol, model with a Python tool, effort max.
- In FrontierMath (tiers 1 to 3), GPT-5.6 Sol, model with a Python tool, effort max, placed above Claude Fable 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above GPT-5.6 Sol, model with a Python tool, effort max.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-6 Astra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-6 Astra, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-6 Astra, sent straight to the model; in FrontierMath (tier 4), GPT-6 Astra, model with a Python tool, effort high, placed above Claude Fable 5.1, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tier 4), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Opus 4.6, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.6, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.6, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In FrontierMath (tier 4), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.2, sent straight to the model, effort high.
- In LiveBench, GPT-5.4, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.4, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
- In FrontierMath (tier 4), GPT-5.4, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.7, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.7, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In FrontierMath (tiers 1 to 3), Claude Opus 4.8, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In FrontierMath (tier 4), Claude Opus 4.8, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.8, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.8, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.8, model with a Python tool, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.8, model with a Python tool, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in OTIS Mock AIME, Claude Opus 5, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above Claude Opus 5, model with a Python tool, effort max.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
- In FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above Claude Opus 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.6 Terra, model with a Python tool, effort max.
- In FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.6 Terra, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
- In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash-Lite, model with a Python tool, effort high, placed above Claude Sonnet 4.5, model with a Python tool, thinking budget 32K; in FrontierMath (tier 4), Claude Sonnet 4.5, model with a Python tool, thinking budget 32K, placed above Gemini 3.5 Flash-Lite, model with a Python tool, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash-Lite, model with a Python tool, effort high, placed above Claude Sonnet 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Sonnet 4.5, sent straight to the model, thinking budget 32K, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 4.6, sent straight to the model; in OTIS Mock AIME, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Sonnet 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort high.
- In FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
- In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in MathArena ArXivMath (June 2026), Gemini 3.1 Pro (preview), sent straight to the model, effort not stated, placed above Gemini 3.5 Flash, sent straight to the model, effort not stated.
- In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort default.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
- In OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low; in MathArena ArXivMath (June 2026), GPT-5.5, sent straight to the model, effort xhigh, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort not stated.
- In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
- In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
- In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort default.
- In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in MathArena ArXivMath (June 2026), GPT-5.5, sent straight to the model, effort xhigh, placed above Gemini 3.5 Flash, sent straight to the model, effort not stated.
- In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high.
- In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high.
- In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.6 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.6 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In FrontierMath (tier 4), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.7 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.7 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.7 Flash, model with a Python tool, effort high.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.7 Flash, model with a Python tool, effort high.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above GPT-5.2, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above GPT-5.2, model with a Python tool, effort xhigh; in FrontierMath (tier 4), GPT-5.2, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tier 4), GPT-5.2, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort high.
- In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.4, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
- In FrontierMath (tier 4), GPT-5.4, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in FrontierMath (tier 4), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh.
- In FrontierMath (tier 4), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5 mini, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5 mini, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.2, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.2, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.2, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.4, sent straight to the model, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in FrontierMath (tier 4), GPT-5.4 Nano, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh.
- In FrontierMath (tier 4), GPT-5.4 Nano, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.4 Nano, sent straight to the model, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Luna, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Luna, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh; in FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Terra, model with a Python tool, effort max.
- In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Terra, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
Quality and cost
In MathArena ArXivMath (June 2026) only. Cost as published by MathArena (ETH Zurich SRI Lab and INSAIT). Your costs will differ.
Most accurate, and the cheapest within 1 point of the best: Claude Fable 5 Sent straight to the model, effort max 85.4% · $3.79 per problem, per run
Best at each price
- 1Gemini 3.5 Flash Sent straight to the model, effort not stated 52.1% · $0.28 per problem, per run
- 2Gemini 3.1 Pro (preview) Sent straight to the model, effort not stated 66.7% · $0.37 per problem, per run
- 3GPT-5.5 Sent straight to the model, effort xhigh 83.6% · $1.82 per problem, per run
- 4Claude Fable 5 Sent straight to the model, effort max 85.4% · $3.79 per problem, per runbest quality
Findings for this category aren't written yet.
LiveBench
best 97.1%LiveBenchNewest result 2026-06-25reliability not measuredcaveat
A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.
Results
All Sent straight to the model · Run date not published; posted 2026-06-25
Each model's best setting in this test
- Claude Opus 5.597.1%
- Claude Fable 5.197.0%
- GPT-6 Astra96.8%
- GPT-6 Sol96.4%
- GPT-5.6 Sol96.2%
- Claude Fable 596.0%
- GPT-5.595.9%
- Claude Opus 595.7%
Show all 29 results (21 not shown above) from LiveBench
- Claude Opus 5.597.1%
- Claude Fable 5.197.0%
- GPT-6 Astra96.8%
- Claude Opus 5.596.8%
- GPT-6 Sol96.4%
- GPT-5.6 Sol96.2%
- Claude Fable 596.0%
- GPT-5.595.9%
- Claude Opus 595.7%
- GPT-5.6 Terra94.9%
- Claude Opus 4.894.3%
- GPT-5.494.2%
- Gemini 3.7 Flash93.5%
- GPT-5.293.2%
- Claude Sonnet 592.9%
- Claude Opus 4.792.8%
- Gemini 3.8 Flash91.6%
- Gemini 3.1 Pro (preview)91.0%
- GPT-5.4 Nano91.0%
- Claude Opus 4.590.4%
- Claude Opus 4.689.3%
- GPT-6 Luna89.1%
- GPT-5.2 Codex88.8%
- Gemini 3.5 Flash88.2%
- GPT-5.6 Luna87.2%
- Claude Sonnet 4.687.0%
- Gemini 3.6 Flash86.4%
- GPT-5.4 Mini78.5%
- Gemini 3.5 Flash-Lite73.7%
Watch out
LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.
Reliability not measured by this source.
LiveBench is funded by Abacus.AI.
See LiveBench's resultsNot tested here: 6 models
Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1.
More about this test: LiveBench
What's in the test
Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.
Technical details for LiveBench
Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
- Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
- Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
Source version: LiveBench release 2026-06-25
Results posted 2026-06-25
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-27.
Sources for this summary:
FrontierMath (tiers 1 to 3)
best 93.7%Epoch AI295 problemsNewest result 2026-09-02reliability not measuredcaveat
Original, very hard math problems written and checked by expert mathematicians, from advanced undergraduate to early research level. Epoch AI runs the test itself, and the model can write and run Python code while it works.
Results
Each model's best setting in this test
- GPT-6 Astra Model with a Python tool, effort max, 2026-08-3093.7%
- Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0190.2%
- GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0989.1%
- Claude Fable 5 Model with a Python tool, effort max, 2026-06-0987.0%
- GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0986.0%
- Claude Opus 5 Model with a Python tool, effort max, 2026-07-2485.6%
- GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1185.3%
- GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0982.1%
Show all 34 results (26 not shown above) from FrontierMath (tiers 1 to 3)
- GPT-6 Astra Model with a Python tool, effort max, 2026-08-3093.7%
- Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0190.2%
- GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0989.1%
- Claude Fable 5 Model with a Python tool, effort max, 2026-06-0987.0%
- GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0986.0%
- Claude Opus 5 Model with a Python tool, effort max, 2026-07-2485.6%
- GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1185.3%
- GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0982.1%
- Claude Opus 4.8 Model with a Python tool, effort max, 2026-06-1080.0%
- GPT-5.4 Model with a Python tool, effort xhigh, 2026-06-1178.6%
- Gemini 3.7 Flash Model with a Python tool, effort high, 2026-08-1471.6%
- Claude Opus 4.7 Model with a Python tool, effort max, 2026-06-1070.2%
- Gemini 3.8 Flash Model with a Python tool, effort high, 2026-09-0268.4%
- GPT-5.2 Model with a Python tool, effort xhigh, 2026-06-1167.4%
- Claude Opus 4.6 Model with a Python tool, effort max, 2026-06-1166.0%
- Claude Sonnet 5 Model with a Python tool, effort max, 2026-06-3065.6%
- Gemini 3.5 Flash Model with a Python tool, effort high, 2026-06-1062.8%
- Gemini 3.1 Pro (preview) Model with a Python tool, effort default, 2026-06-1159.6%
- Gemini 3.6 Flash Model with a Python tool, effort high, 2026-08-0258.9%
- Gemini 3 Flash (preview) Model with a Python tool, effort default, 2026-06-1151.2%
- GPT-5.4 Mini Model with a Python tool, effort xhigh, 2026-06-1251.2%
- GPT-5 mini Model with a Python tool, effort high, 2026-06-1246.7%
- GPT-5.4 Nano Model with a Python tool, effort high, 2026-06-1244.9%
- GPT-5.6 Luna Model with a Python tool, effort low, 2026-08-2941.4%
- GPT-5.6 Luna Model with a Python tool, effort none, 2026-08-2939.6%
- Claude Opus 4.5 Model with a Python tool, thinking budget 32K, 2026-06-1134.4%
- Gemini 3.5 Flash-Lite Model with a Python tool, effort high, 2026-08-0226.0%
- GPT-5.4 Mini Model with a Python tool, effort low, 2026-08-2824.6%
- Claude Sonnet 4.5 Model with a Python tool, thinking budget 32K, 2026-06-1123.9%
- GPT-5.4 Nano Model with a Python tool, effort low, 2026-08-2820.4%
- GPT-5 mini Model with a Python tool, effort low, 2026-08-2718.2%
- GPT-5.4 Mini Model with a Python tool, effort none, 2026-08-2817.2%
- GPT-5 mini Model with a Python tool, effort minimal, 2026-08-276.0%
- GPT-5.4 Nano Model with a Python tool, effort none, 2026-08-284.6%
Watch out
Epoch released version 2 on 2026-06-12 after fixing errors in 42% of FrontierMath problems. Scores here are on version 2. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability not measured by this source.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute. Epoch AI says FrontierMath was developed with funding from OpenAI, which has exclusive access to a subset of the benchmark.
See Epoch AI's resultsNot tested here: 8 models
Claude Haiku 4.5, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-5.1, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.
More about this test: FrontierMath (tiers 1 to 3)
What's in the test
295 problems covering most major branches of modern mathematics. A typical problem takes a researcher in that field several hours. The model submits a Python function that returns its answer, which is checked automatically. Scores are on Epoch's private problems.
295 problems
Technical details for FrontierMath (tiers 1 to 3)
Accuracy: The share of the 295 problems the model solved, in Epoch AI's own runs.
- GPT-6 Astra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Fable 5.1, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.6 Sol, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Fable 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.6 Terra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Opus 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.5, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- GPT-5.6 Luna, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Opus 4.8, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.4, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- Gemini 3.7 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Claude Opus 4.7, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Gemini 3.8 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.2, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- Claude Opus 4.6, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Sonnet 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Gemini 3.5 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Gemini 3.1 Pro (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
- Gemini 3.6 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Gemini 3 Flash (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
- GPT-5.4 Mini, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- GPT-5 mini, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.4 Nano, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.6 Luna, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
- GPT-5.6 Luna, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
- Claude Opus 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
- Gemini 3.5 Flash-Lite, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.4 Mini, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
- Claude Sonnet 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
- GPT-5.4 Nano, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
- GPT-5 mini, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
- GPT-5.4 Mini, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
- GPT-5 mini, model with a Python tool, effort minimal: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: minimal.
- GPT-5.4 Nano, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2026-06-09 to 2026-09-02
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
FrontierMath (tier 4)
best 97.6%Epoch AI43 problemsNewest result 2026-09-02reliability not measuredcaveat
The hardest tier of FrontierMath: exceptionally difficult research-level math problems written and checked by expert mathematicians. Epoch AI runs the test itself, and the model can write and run Python code while it works.
Results
Each model's best setting in this test
- GPT-6 Astra Model with a Python tool, effort high, 2026-08-3097.6%
- Claude Fable 5 Model with a Python tool, effort max, 2026-06-0990.2%
- Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0187.8%
- GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0982.9%
- Claude Opus 5 Model with a Python tool, effort max, 2026-07-2473.2%
- GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1172.5%
- GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0970.7%
- GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0961.0%
Show all 31 results (23 not shown above) from FrontierMath (tier 4)
- GPT-6 Astra Model with a Python tool, effort high, 2026-08-3097.6%
- GPT-6 Astra Model with a Python tool, effort max, 2026-08-3097.6%
- GPT-6 Astra Model with a Python tool, effort xhigh, 2026-08-3097.6%
- GPT-6 Astra Model with a Python tool, effort medium, 2026-08-3097.6%
- Claude Fable 5 Model with a Python tool, effort max, 2026-06-0990.2%
- Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0187.8%
- GPT-6 Astra Model with a Python tool, effort low, 2026-08-3087.8%
- GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0982.9%
- GPT-6 Astra Model with a Python tool, effort none, 2026-08-3082.9%
- Claude Opus 5 Model with a Python tool, effort max, 2026-07-2473.2%
- GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1172.5%
- GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0970.7%
- GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0961.0%
- Claude Opus 4.8 Model with a Python tool, effort max, 2026-06-1056.1%
- GPT-5.4 Model with a Python tool, effort xhigh, 2026-06-1149.0%
- Gemini 3.7 Flash Model with a Python tool, effort high, 2026-08-1436.6%
- Claude Opus 4.7 Model with a Python tool, effort max, 2026-06-1031.7%
- GPT-5.2 Model with a Python tool, effort xhigh, 2026-06-1131.7%
- Claude Sonnet 5 Model with a Python tool, effort max, 2026-06-3029.3%
- Claude Opus 4.6 Model with a Python tool, effort max, 2026-06-1126.8%
- Gemini 3.1 Pro (preview) Model with a Python tool, effort default, 2026-06-1126.8%
- Gemini 3.5 Flash Model with a Python tool, effort high, 2026-06-1026.8%
- Gemini 3.6 Flash Model with a Python tool, effort high, 2026-08-0222.0%
- Gemini 3.8 Flash Model with a Python tool, effort high, 2026-09-0222.0%
- Gemini 3 Flash (preview) Model with a Python tool, effort default, 2026-06-1117.1%
- GPT-5 mini Model with a Python tool, effort high, 2026-06-1212.2%
- GPT-5.4 Nano Model with a Python tool, effort high, 2026-06-1212.2%
- GPT-5.4 Mini Model with a Python tool, effort xhigh, 2026-06-129.8%
- Claude Opus 4.5 Model with a Python tool, thinking budget 32K, 2026-06-114.9%
- Claude Sonnet 4.5 Model with a Python tool, thinking budget 32K, 2026-06-112.4%
- Gemini 3.5 Flash-Lite Model with a Python tool, effort high, 2026-08-020.0%
Watch out
With only 43 problems, each one is worth more than 2 points. Epoch released version 2 on 2026-06-12 after fixing errors in 42% of FrontierMath problems. Scores here are on version 2. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability not measured by this source.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute. Epoch AI says FrontierMath was developed with funding from OpenAI, which has exclusive access to a subset of the benchmark.
See Epoch AI's resultsNot tested here: 8 models
Claude Haiku 4.5, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-5.1, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.
More about this test: FrontierMath (tier 4)
What's in the test
43 problems. The hardest can take a researcher in that field several days. The model submits a Python function that returns its answer, which is checked automatically. Scores are on Epoch's private problems.
43 problems
Technical details for FrontierMath (tier 4)
Accuracy: The share of the 43 problems the model solved, in Epoch AI's own runs.
- GPT-6 Astra, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-6 Astra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-6 Astra, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- GPT-6 Astra, model with a Python tool, effort medium: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: medium.
- Claude Fable 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Fable 5.1, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-6 Astra, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
- GPT-5.6 Sol, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-6 Astra, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
- Claude Opus 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.5, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- GPT-5.6 Terra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.6 Luna, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Opus 4.8, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.4, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- Gemini 3.7 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Claude Opus 4.7, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- GPT-5.2, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- Claude Sonnet 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Claude Opus 4.6, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
- Gemini 3.1 Pro (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
- Gemini 3.5 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Gemini 3.6 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Gemini 3.8 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- Gemini 3 Flash (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
- GPT-5 mini, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.4 Nano, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
- GPT-5.4 Mini, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
- Claude Opus 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
- Claude Sonnet 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
- Gemini 3.5 Flash-Lite, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2026-06-09 to 2026-09-02
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
OTIS Mock AIME
best 100.0%Epoch AI45 problemsNewest result 2026-09-02repeats runs, method under the sourcecaveat
Competition-style math problems from the OTIS Mock AIME exams, written by students in the OTIS olympiad training program. Epoch AI runs the test itself.
Results
Each model's best setting in this test
- Claude Fable 5 Sent straight to the model, effort high, 2026-08-06100.0%
- Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-01100.0%
- GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-09100.0%
- GPT-6 Astra Sent straight to the model, effort max, 2026-08-30100.0%
- GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0999.7%
- Claude Opus 5 Sent straight to the model, effort max, 2026-07-2498.9%
- Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0298.9%
- Claude Opus 4.8 Sent straight to the model, effort max, 2026-06-0798.3%
- GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0998.3%
Show all 80 results (71 not shown above) from OTIS Mock AIME
- Claude Fable 5 Sent straight to the model, effort high, 2026-08-06100.0%
- Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-01100.0%
- GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-09100.0%
- GPT-6 Astra Sent straight to the model, effort max, 2026-08-30100.0%
- Claude Fable 5 Sent straight to the model, effort max, 2026-06-1099.7%
- GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0999.7%
- Claude Opus 5 Sent straight to the model, effort max, 2026-07-2498.9%
- Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0298.9%
- Claude Opus 4.8 Sent straight to the model, effort max, 2026-06-0798.3%
- GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0998.3%
- Claude Opus 4.7 Sent straight to the model, effort xhigh, 2026-04-1797.8%
- Claude Fable 5 Sent straight to the model, effort low, 2026-08-0697.8%
- Claude Opus 4.8 Sent straight to the model, effort low, 2026-08-0697.8%
- Claude Opus 5 Sent straight to the model, effort default, 2026-08-0697.8%
- GPT-5.4 Sent straight to the model, effort high, 2026-07-1597.8%
- Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1497.2%
- GPT-5.2 Sent straight to the model, effort high, 2025-12-1196.1%
- GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1396.1%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-2095.6%
- Gemini 3 Flash (preview) Sent straight to the model, effort high, 2026-08-0695.6%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0695.6%
- Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2595.6%
- GPT-5.4 Sent straight to the model, effort medium, 2026-07-1595.6%
- GPT-5.6 Sol Sent straight to the model, effort low, 2026-08-0795.6%
- GPT-5.4 Sent straight to the model, effort xhigh, 2026-03-0695.3%
- Claude Sonnet 5 Sent straight to the model, effort xhigh, 2026-07-0194.7%
- Claude Opus 4.6 Sent straight to the model, thinking budget 64K, 2026-02-0694.4%
- Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0294.2%
- GPT-5.2 Sent straight to the model, effort medium, 2025-12-1193.9%
- Claude Opus 5 Sent straight to the model, effort low, 2026-08-0693.3%
- Claude Opus 4.6 Sent straight to the model, thinking budget 32K, 2026-02-0693.1%
- Gemini 3 Flash (preview) Sent straight to the model, effort default, 2025-12-1792.8%
- Gemini 3 Pro (preview) Sent straight to the model, effort default, 2025-11-1991.4%
- Claude Opus 4.6 Sent straight to the model, effort max, 2026-08-0691.1%
- Gemini 3.5 Flash Sent straight to the model, effort low, 2026-08-0688.9%
- GPT-5.4 Mini Sent straight to the model, effort xhigh, 2026-08-0788.9%
- GPT-5.6 Terra Sent straight to the model, effort low, 2026-08-0788.9%
- GPT-5.1 Sent straight to the model, effort high, 2025-11-1388.6%
- GPT-5.4 Nano Sent straight to the model, effort high, 2026-04-1487.8%
- GPT-5.4 Mini Sent straight to the model, effort high, 2026-04-1587.2%
- Claude Opus 4.7 Sent straight to the model, effort max, 2026-08-0686.7%
- GPT-5 mini Sent straight to the model, effort high, 2025-10-3086.7%
- Claude Opus 4.5 Sent straight to the model, thinking budget 32K, 2025-11-2486.1%
- Claude Sonnet 4.6 Sent straight to the model, thinking budget 32K, 2026-02-2085.8%
- GPT-5.1 Sent straight to the model, effort medium, 2025-11-1785.6%
- Claude Opus 4.8 Sent straight to the model, effort none, 2026-08-0684.4%
- GPT-5.4 Sent straight to the model, effort low, 2026-07-1584.4%
- GPT-5.5 Sent straight to the model, effort low, 2026-08-0784.4%
- Claude Sonnet 4.6 Sent straight to the model, effort medium, 2026-07-1382.2%
- Gemini 3.6 Flash Sent straight to the model, effort low, 2026-08-0782.2%
- Claude Opus 4.5 Sent straight to the model, thinking budget 16K, 2025-11-2481.7%
- Claude Sonnet 5 Sent straight to the model, effort max, 2026-08-0680.0%
- Gemini 3.5 Flash Sent straight to the model, effort minimal, 2026-07-1580.0%
- Gemini 3.6 Flash Sent straight to the model, effort minimal, 2026-08-0780.0%
- GPT-5.2 Sent straight to the model, effort low, 2025-12-1178.9%
- Claude Sonnet 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2177.8%
- Claude Sonnet 4.5 Sent straight to the model, thinking budget 59K, 2025-10-2877.8%
- Claude Sonnet 4.6 Sent straight to the model, effort high, 2026-07-1375.6%
- Claude Sonnet 4.5 Sent straight to the model, thinking budget 16K, 2025-10-2871.1%
- Claude Sonnet 4.6 Sent straight to the model, effort max, 2026-08-0671.1%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort high, 2026-08-0671.1%
- GPT-5.4 Nano Sent straight to the model, effort low, 2026-08-0768.9%
- GPT-5.6 Sol Sent straight to the model, effort none, 2026-08-0768.9%
- Claude Haiku 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2266.7%
- GPT-5.6 Luna Sent straight to the model, effort low, 2026-08-0766.7%
- GPT-5.1 Sent straight to the model, effort low, 2025-11-2563.9%
- GPT-5.2 Sent straight to the model, effort none, 2026-07-1362.2%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort low, 2026-08-0660.0%
- GPT-5.4 Sent straight to the model, effort none, 2026-07-1557.8%
- GPT-5.5 Sent straight to the model, effort none, 2026-08-0757.8%
- GPT-5 mini Sent straight to the model, effort minimal, 2026-08-0755.6%
- GPT-5.6 Terra Sent straight to the model, effort none, 2026-08-0753.3%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort minimal, 2026-08-0651.1%
- Claude Opus 4.5 Sent straight to the model, effort default, 2025-11-2448.1%
- GPT-5.4 Nano Sent straight to the model, effort none, 2026-08-0746.7%
- GPT-5.6 Luna Sent straight to the model, effort none, 2026-08-0740.0%
- GPT-5.1 Sent straight to the model, effort none, 2026-08-0737.8%
- Claude Haiku 4.5 Sent straight to the model, effort default, 2025-10-1635.8%
- Claude Sonnet 4.5 Sent straight to the model, effort default, 2025-09-2935.6%
- GPT-5.4 Mini Sent straight to the model, effort none, 2026-08-0726.7%
Watch out
With 45 problems, each one is worth more than 2 points. Three problems include a picture; Epoch leaves the picture out and gives every model the code that draws it instead. In September 2025 Epoch changed how it reads the final answer, after its earlier grader marked some blank answers correct. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability: Epoch AI says it runs most models 16 times on this test and reports the average score. How many runs each model got is not in its download, and a few scores do not fit 16 runs.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.
See Epoch AI's resultsNot tested here: 4 models
Claude Opus 5.5, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.
More about this test: OTIS Mock AIME
What's in the test
45 problems from three mock exams (2024, 2025 I and 2025 II), 15 from each. Every answer is a whole number from 0 to 999. Epoch rates them harder than MATH Level 5 but easier than FrontierMath.
45 problems
Older results, more than a year old (1) from OTIS Mock AIME
- GPT-5 mini Sent straight to the model, effort medium, 2025-08-0778.3%
Technical details for OTIS Mock AIME
Accuracy: The share of the 45 problems the model answered correctly, in Epoch AI's own runs.
- Claude Fable 5, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Fable 5.1, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.6 Sol, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-6 Astra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Fable 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.6 Terra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Opus 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.8 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Opus 4.8, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.6 Luna, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Opus 4.7, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Claude Fable 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.8, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.4, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.7 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Gemini 3.1 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Gemini 3 Flash (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.1 Pro (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.5 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- GPT-5.6 Sol, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Claude Sonnet 5, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Claude Opus 4.6, sent straight to the model, thinking budget 64K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 64K tokens.
- Gemini 3.6 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- Claude Opus 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Gemini 3 Flash (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Gemini 3 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Claude Opus 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.5 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4 Mini, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- GPT-5.6 Terra, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.1, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4 Nano, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4 Mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Opus 4.7, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5 mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Opus 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Claude Sonnet 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- GPT-5.1, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- Claude Opus 4.8, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.4, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Sonnet 4.6, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- Gemini 3.6 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
- Claude Sonnet 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.5 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- Gemini 3.6 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- GPT-5.2, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Sonnet 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Claude Sonnet 4.5, sent straight to the model, thinking budget 59K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 59K tokens.
- Claude Sonnet 4.6, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Sonnet 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
- Claude Sonnet 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4 Nano, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.6 Sol, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Haiku 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- GPT-5.6 Luna, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.1, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.2, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.5, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5 mini, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- GPT-5.6 Terra, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- Claude Opus 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.4 Nano, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.6 Luna, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.1, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Haiku 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Claude Sonnet 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.4 Mini, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2025-09-29 to 2026-09-02
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
MathArena ArXivMath (June 2026)
best 85.4%MathArena (ETH Zurich SRI Lab and INSAIT)48 problemsNewest result 2026-09-07reliability not measuredcaveat
MathArena, from ETH Zurich, tests AI on new maths problems soon after they appear, so models are less likely to have seen them. ArXivMath takes its problems from research papers posted on arXiv each month; each problem has one correct final answer.
Results
Run date not published; posted 2026-09-07
- Claude Fable 5 Sent straight to the model, effort max85.4%
- GPT-5.5 Sent straight to the model, effort xhigh83.6%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort not stated66.7%
- Gemini 3.5 Flash Sent straight to the model, effort not stated52.1%
Watch out
MathArena does not publish when it ran each model. The date shown is when MathArena published that model's answers. The number of runs differs by model and is shown with each result; MathArena's site says it runs each model 4 times, but its published answers show 3 to 7 runs per problem. MathArena marks 1 of the models shown as released after these problems were published, so it may have seen the source papers in training: Claude-Fable-5 (max). Cost is MathArena's estimate for one run on one problem at the model maker's prices. With 48 problems, each one is worth about 2 points. Results MathArena has not published answers for, or whose published answers do not match its site, are not shown. Nor are results whose settings MathArena published only after their answers, since the settings used for those runs cannot be confirmed. Data: MathArena, ETH Zurich, matharena.ai, CC BY-SA 4.0, shown with MathArena's permission.
Reliability not measured by this source.
MathArena is run by the SRI Lab at ETH Zurich and INSAIT, which lists Google and DeepMind among its supporters. Google makes Gemini, one of the model families shown. MathArena does not make any of the models it tests.
See MathArena (ETH Zurich SRI Lab and INSAIT)'s resultsNot tested here: 30 models
Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5 mini, GPT-5.1, GPT-5.2, GPT-5.2 Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
More about this test: MathArena ArXivMath (June 2026)
What's in the test
48 research-level problems taken from maths papers submitted to arXiv in June 2026. The model must give the final answer, which is checked against the answer from the paper. MathArena publishes a new set each month; only the June 2026 set is shown, because each month is a different test.
48 problems
- Research-level maths
- Final-answer problems
Technical details for MathArena ArXivMath (June 2026)
Accuracy: The share of the 48 research-level maths problems the model answered correctly, averaged over MathArena's runs of each problem.
- Claude Fable 5, sent straight to the model, effort max: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort: max.
- GPT-5.5, sent straight to the model, effort xhigh: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort: xhigh.
- Gemini 3.1 Pro (preview), sent straight to the model, effort not stated: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort not stated.
- Gemini 3.5 Flash, sent straight to the model, effort not stated: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort not stated.
Source version: MathArena ArXivMath June 2026 answers at commit f782eef, downloaded 2026-09-30
Results posted 2026-09-07
Run dates not published by the tester; the dates are when results were posted.
License: CC BY-SA 4.0 (MathArena's ArXivMath data); display permitted by MathArena in writing, 2026-09-28. Checked 2026-09-30.
No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano, GPT-5.1 Codex.
What this doesn't tell you
- Math tests use set problems with known answers. Problems in your own work may be less tidy.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- LiveBench: Reliability not measured by this source.
- FrontierMath (tiers 1 to 3): Reliability not measured by this source.
- FrontierMath (tier 4): Reliability not measured by this source.
- MathArena ArXivMath (June 2026): Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use