Assurance

Mathematics

Solving mathematical problems, such as working out quantities in a word problem, checking a calculation, or finding the steps of a proof. A good result reaches the correct answer, not just a plausible one.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Direct sent straight to the modelAgent works through the task with tools, as the words after it sayFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, earned repeat runs, some or all models: Results checked by repeat runs, for some or all models, earned

Vendor claims have not been collected for any category yet.

LB 97.1%LB 97.1%FM 93.7%FM 93.7%FM4 97.6%FM4 97.6%AIME 100.0%AIME 100.0%AXM 85.4%AXM 85.4%
ⓘ More about this category
  • A source here repeats runs for some or all models; its method is under the source. Per-model figures not shown.
  • 5 sources, 34 models with figures
  • Sources last checked 2026-09-30
  • Ring 4: the tester says it repeats its runs; its method is under the source.
  • LB = LiveBench, FM = FrontierMath (tiers 1 to 3), FM4 = FrontierMath (tier 4), AIME = OTIS Mock AIME, AXM = MathArena ArXivMath (June 2026)

Who does well

TestBestBest scoreNextNext score
LiveBenchClaude Opus 5.5 Sent straight to the modelroughly tied, 5 setupsAlso roughly tied: GPT-6 Astra Sent straight to the model 96.8%; GPT-6 Sol Sent straight to the model 96.4%; GPT-5.6 Sol Sent straight to the model 96.2%97.1%Claude Fable 5.1 Sent straight to the model97.0%
FrontierMath (tiers 1 to 3)GPT-6 Astra Model with a Python tool, effort max93.7%Claude Fable 5.1 Model with a Python tool, effort max90.2%
FrontierMath (tier 4)GPT-6 Astra Model with a Python tool, effort high97.6%Claude Fable 5 Model with a Python tool, effort max90.2%
OTIS Mock AIMEClaude Fable 5 Sent straight to the model, effort highroughly tied, 5 setupsAlso roughly tied: GPT-5.6 Sol Sent straight to the model, effort max 100.0%; GPT-6 Astra Sent straight to the model, effort max 100.0%; GPT-5.6 Terra Sent straight to the model, effort max 99.7%100.0%Claude Fable 5.1 Sent straight to the model, effort max100.0%
MathArena ArXivMath (June 2026)Claude Fable 5 Sent straight to the model, effort max85.4%GPT-5.5 Sent straight to the model, effort xhigh83.6%
Place within each test. Places are never added up across tests.
ModelLiveBenchFrontierMath (tiers 1 to 3)FrontierMath (tier 4)OTIS Mock AIMEMathArena ArXivMath (June 2026)
Claude Fable 56 of 28Sent straight to the model4 of 26Model with a Python tool, effort max2 of 26Model with a Python tool, effort max1 of 30Sent straight to the model, effort high1 of 4Sent straight to the model, effort max
Claude Fable 5.12 of 28Sent straight to the model2 of 26Model with a Python tool, effort max3 of 26Model with a Python tool, effort max1 of 30Sent straight to the model, effort max
Claude Opus 4.519 of 28Sent straight to the model24 of 26Model with a Python tool, thinking budget 32K24 of 26Model with a Python tool, thinking budget 32K25 of 30Sent straight to the model, thinking budget 32K
Claude Opus 4.620 of 28Sent straight to the model15 of 26Model with a Python tool, effort max15 of 26Model with a Python tool, effort max18 of 30Sent straight to the model, thinking budget 64K
Claude Opus 4.715 of 28Sent straight to the model12 of 26Model with a Python tool, effort max12 of 26Model with a Python tool, effort max10 of 30Sent straight to the model, effort xhigh
Claude Opus 4.810 of 28Sent straight to the model9 of 26Model with a Python tool, effort max9 of 26Model with a Python tool, effort max8 of 30Sent straight to the model, effort max
Claude Opus 58 of 28Sent straight to the model6 of 26Model with a Python tool, effort max5 of 26Model with a Python tool, effort max6 of 30Sent straight to the model, effort max
Claude Sonnet 4.526 of 26Model with a Python tool, thinking budget 32K25 of 26Model with a Python tool, thinking budget 32K28 of 30Sent straight to the model, thinking budget 32K
Claude Sonnet 4.625 of 28Sent straight to the model26 of 30Sent straight to the model, thinking budget 32K
Claude Sonnet 514 of 28Sent straight to the model16 of 26Model with a Python tool, effort max14 of 26Model with a Python tool, effort max17 of 30Sent straight to the model, effort xhigh
Gemini 3 Flash (preview)20 of 26Model with a Python tool, effort default20 of 26Model with a Python tool, effort default15 of 30Sent straight to the model, effort high
Gemini 3.1 Pro (preview)17 of 28Sent straight to the model18 of 26Model with a Python tool, effort default15 of 26Model with a Python tool, effort default14 of 30Sent straight to the model, effort default3 of 4Sent straight to the model, effort not stated
Gemini 3.5 Flash23 of 28Sent straight to the model17 of 26Model with a Python tool, effort high15 of 26Model with a Python tool, effort high15 of 30Sent straight to the model, effort high4 of 4Sent straight to the model, effort not stated
Gemini 3.5 Flash-Lite28 of 28Sent straight to the model25 of 26Model with a Python tool, effort high26 of 26Model with a Python tool, effort high29 of 30Sent straight to the model, effort high
Gemini 3.6 Flash26 of 28Sent straight to the model19 of 26Model with a Python tool, effort high18 of 26Model with a Python tool, effort high19 of 30Sent straight to the model, effort high
Gemini 3.7 Flash12 of 28Sent straight to the model11 of 26Model with a Python tool, effort high11 of 26Model with a Python tool, effort high12 of 30Sent straight to the model, effort high
Gemini 3.8 Flash16 of 28Sent straight to the model13 of 26Model with a Python tool, effort high18 of 26Model with a Python tool, effort high6 of 30Sent straight to the model, effort high
GPT-5 mini22 of 26Model with a Python tool, effort high21 of 26Model with a Python tool, effort high24 of 30Sent straight to the model, effort high
GPT-5.213 of 28Sent straight to the model14 of 26Model with a Python tool, effort xhigh13 of 26Model with a Python tool, effort xhigh13 of 30Sent straight to the model, effort high
GPT-5.411 of 28Sent straight to the model10 of 26Model with a Python tool, effort xhigh10 of 26Model with a Python tool, effort xhigh11 of 30Sent straight to the model, effort high
GPT-5.4 Mini27 of 28Sent straight to the model20 of 26Model with a Python tool, effort xhigh23 of 26Model with a Python tool, effort xhigh21 of 30Sent straight to the model, effort xhigh
GPT-5.4 Nano18 of 28Sent straight to the model23 of 26Model with a Python tool, effort high21 of 26Model with a Python tool, effort high23 of 30Sent straight to the model, effort high
GPT-5.57 of 28Sent straight to the model7 of 26Model with a Python tool, effort xhigh6 of 26Model with a Python tool, effort xhigh27 of 30Sent straight to the model, effort low2 of 4Sent straight to the model, effort xhigh
GPT-5.6 Luna24 of 28Sent straight to the model8 of 26Model with a Python tool, effort max8 of 26Model with a Python tool, effort max8 of 30Sent straight to the model, effort max
GPT-5.6 Sol5 of 28Sent straight to the model3 of 26Model with a Python tool, effort max4 of 26Model with a Python tool, effort max1 of 30Sent straight to the model, effort max
GPT-5.6 Terra9 of 28Sent straight to the model5 of 26Model with a Python tool, effort max7 of 26Model with a Python tool, effort max5 of 30Sent straight to the model, effort max
GPT-6 Astra3 of 28Sent straight to the model1 of 26Model with a Python tool, effort max1 of 26Model with a Python tool, effort high1 of 30Sent straight to the model, effort max

The tests here put 200 pairs of models in a different order.

Show all 200 pairs placed in a different order
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
  • In FrontierMath (tiers 1 to 3), Claude Fable 5.1, model with a Python tool, effort max, placed above Claude Fable 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
  • In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above GPT-5.6 Sol, model with a Python tool, effort max.
  • In FrontierMath (tiers 1 to 3), GPT-5.6 Sol, model with a Python tool, effort max, placed above Claude Fable 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Fable 5, model with a Python tool, effort max, placed above GPT-5.6 Sol, model with a Python tool, effort max.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-6 Astra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-6 Astra, model with a Python tool, effort max, placed above Claude Fable 5.1, model with a Python tool, effort max.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-6 Astra, sent straight to the model; in FrontierMath (tier 4), GPT-6 Astra, model with a Python tool, effort high, placed above Claude Fable 5.1, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tier 4), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Opus 4.5, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.5, model with a Python tool, thinking budget 32K.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Opus 4.6, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.6, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.6, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.6, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.6, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.6, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.6, sent straight to the model, thinking budget 64K, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.6, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 64K.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In FrontierMath (tier 4), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 4.7, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in FrontierMath (tier 4), Claude Opus 4.7, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.2, sent straight to the model, effort high.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.4, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
  • In FrontierMath (tier 4), GPT-5.4, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.7, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.7, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.7, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In FrontierMath (tiers 1 to 3), Claude Opus 4.8, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In FrontierMath (tier 4), Claude Opus 4.8, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.8, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Opus 4.8, model with a Python tool, effort max; in OTIS Mock AIME, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.8, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Opus 4.8, model with a Python tool, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in FrontierMath (tiers 1 to 3), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in OTIS Mock AIME, Claude Opus 5, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above Claude Opus 5, model with a Python tool, effort max.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
  • In FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above Claude Opus 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.6 Terra, model with a Python tool, effort max.
  • In FrontierMath (tier 4), Claude Opus 5, model with a Python tool, effort max, placed above GPT-5.6 Terra, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
  • In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash-Lite, model with a Python tool, effort high, placed above Claude Sonnet 4.5, model with a Python tool, thinking budget 32K; in FrontierMath (tier 4), Claude Sonnet 4.5, model with a Python tool, thinking budget 32K, placed above Gemini 3.5 Flash-Lite, model with a Python tool, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash-Lite, model with a Python tool, effort high, placed above Claude Sonnet 4.5, model with a Python tool, thinking budget 32K; in OTIS Mock AIME, Claude Sonnet 4.5, sent straight to the model, thinking budget 32K, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 4.6, sent straight to the model; in OTIS Mock AIME, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tiers 1 to 3), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Claude Sonnet 5, model with a Python tool, effort max; in FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tier 4), Claude Sonnet 5, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Claude Sonnet 5, model with a Python tool, effort max; in OTIS Mock AIME, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Claude Sonnet 5, model with a Python tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort high.
  • In FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3 Flash (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
  • In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in MathArena ArXivMath (June 2026), Gemini 3.1 Pro (preview), sent straight to the model, effort not stated, placed above Gemini 3.5 Flash, sent straight to the model, effort not stated.
  • In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tier 4), Gemini 3.1 Pro (preview), model with a Python tool, effort default, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort default.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default; in OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
  • In OTIS Mock AIME, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low; in MathArena ArXivMath (June 2026), GPT-5.5, sent straight to the model, effort xhigh, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort not stated.
  • In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
  • In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.1 Pro (preview), model with a Python tool, effort default.
  • In LiveBench, Gemini 3.1 Pro (preview), sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.1 Pro (preview), sent straight to the model, effort default.
  • In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.5 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.5 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.5 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In OTIS Mock AIME, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in MathArena ArXivMath (June 2026), GPT-5.5, sent straight to the model, effort xhigh, placed above Gemini 3.5 Flash, sent straight to the model, effort not stated.
  • In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high.
  • In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.5 Flash, model with a Python tool, effort high.
  • In LiveBench, Gemini 3.5 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.5 Flash, sent straight to the model, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.6 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in FrontierMath (tier 4), Gemini 3.6 Flash, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.6 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.6 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.6 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In FrontierMath (tier 4), Gemini 3.7 Flash, model with a Python tool, effort high, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.7 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.7 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.7 Flash, model with a Python tool, effort high.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.7 Flash, model with a Python tool, effort high.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above GPT-5.2, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), Gemini 3.8 Flash, model with a Python tool, effort high, placed above GPT-5.2, model with a Python tool, effort xhigh; in FrontierMath (tier 4), GPT-5.2, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tier 4), GPT-5.2, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort high.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.4, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
  • In FrontierMath (tier 4), GPT-5.4, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In LiveBench, Gemini 3.8 Flash, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above Gemini 3.8 Flash, model with a Python tool, effort high; in OTIS Mock AIME, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in FrontierMath (tier 4), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh.
  • In FrontierMath (tier 4), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5 mini, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5 mini, model with a Python tool, effort high, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5 mini, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5 mini, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.2, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.2, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.2, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.2, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.2, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.4, sent straight to the model, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In FrontierMath (tiers 1 to 3), GPT-5.4 Mini, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in FrontierMath (tier 4), GPT-5.4 Nano, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh.
  • In FrontierMath (tier 4), GPT-5.4 Nano, model with a Python tool, effort high, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Mini, model with a Python tool, effort xhigh; in OTIS Mock AIME, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.4 Nano, model with a Python tool, effort high; in OTIS Mock AIME, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in FrontierMath (tier 4), GPT-5.6 Luna, model with a Python tool, effort max, placed above GPT-5.4 Nano, model with a Python tool, effort high.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.4 Nano, sent straight to the model, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Luna, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Luna, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In FrontierMath (tiers 1 to 3), GPT-5.6 Terra, model with a Python tool, effort max, placed above GPT-5.5, model with a Python tool, effort xhigh; in FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Terra, model with a Python tool, effort max.
  • In FrontierMath (tier 4), GPT-5.5, model with a Python tool, effort xhigh, placed above GPT-5.6 Terra, model with a Python tool, effort max; in OTIS Mock AIME, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.

Quality and cost

In MathArena ArXivMath (June 2026) only. Cost as published by MathArena (ETH Zurich SRI Lab and INSAIT). Your costs will differ.

Most accurate, and the cheapest within 1 point of the best: Claude Fable 5 Sent straight to the model, effort max 85.4% · $3.79 per problem, per run

Best at each price

  1. 1Gemini 3.5 Flash Sent straight to the model, effort not stated 52.1% · $0.28 per problem, per run
  2. 2Gemini 3.1 Pro (preview) Sent straight to the model, effort not stated 66.7% · $0.37 per problem, per run
  3. 3GPT-5.5 Sent straight to the model, effort xhigh 83.6% · $1.82 per problem, per run
  4. 4Claude Fable 5 Sent straight to the model, effort max 85.4% · $3.79 per problem, per runbest quality

Findings for this category aren't written yet.

The tests behind this

LiveBench

best 97.1%LiveBenchNewest result 2026-06-25reliability not measuredcaveat

A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.

Results

All Sent straight to the model · Run date not published; posted 2026-06-25

Each model's best setting in this test

  • Claude Opus 5.597.1%
  • Claude Fable 5.197.0%
  • GPT-6 Astra96.8%
  • GPT-6 Sol96.4%
  • GPT-5.6 Sol96.2%
  • Claude Fable 596.0%
  • GPT-5.595.9%
  • Claude Opus 595.7%
Show all 29 results (21 not shown above) from LiveBench
  • Claude Opus 5.597.1%
  • Claude Fable 5.197.0%
  • GPT-6 Astra96.8%
  • Claude Opus 5.596.8%
  • GPT-6 Sol96.4%
  • GPT-5.6 Sol96.2%
  • Claude Fable 596.0%
  • GPT-5.595.9%
  • Claude Opus 595.7%
  • GPT-5.6 Terra94.9%
  • Claude Opus 4.894.3%
  • GPT-5.494.2%
  • Gemini 3.7 Flash93.5%
  • GPT-5.293.2%
  • Claude Sonnet 592.9%
  • Claude Opus 4.792.8%
  • Gemini 3.8 Flash91.6%
  • Gemini 3.1 Pro (preview)91.0%
  • GPT-5.4 Nano91.0%
  • Claude Opus 4.590.4%
  • Claude Opus 4.689.3%
  • GPT-6 Luna89.1%
  • GPT-5.2 Codex88.8%
  • Gemini 3.5 Flash88.2%
  • GPT-5.6 Luna87.2%
  • Claude Sonnet 4.687.0%
  • Gemini 3.6 Flash86.4%
  • GPT-5.4 Mini78.5%
  • Gemini 3.5 Flash-Lite73.7%

Watch out

LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.

Reliability not measured by this source.

LiveBench is funded by Abacus.AI.

See LiveBench's results
Not tested here: 6 models

Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1.

More about this test: LiveBench

What's in the test

Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.

Technical details for LiveBench

Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.

  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
  • Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
  • Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.

Source version: LiveBench release 2026-06-25

Results posted 2026-06-25

Run dates not published by the tester; the dates are when results were posted.

License: Apache License 2.0. Checked 2026-09-27.

FrontierMath (tiers 1 to 3)

best 93.7%Epoch AI295 problemsNewest result 2026-09-02reliability not measuredcaveat

Original, very hard math problems written and checked by expert mathematicians, from advanced undergraduate to early research level. Epoch AI runs the test itself, and the model can write and run Python code while it works.

Results

Each model's best setting in this test

  • GPT-6 Astra Model with a Python tool, effort max, 2026-08-3093.7%
  • Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0190.2%
  • GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0989.1%
  • Claude Fable 5 Model with a Python tool, effort max, 2026-06-0987.0%
  • GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0986.0%
  • Claude Opus 5 Model with a Python tool, effort max, 2026-07-2485.6%
  • GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1185.3%
  • GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0982.1%
Show all 34 results (26 not shown above) from FrontierMath (tiers 1 to 3)
  • GPT-6 Astra Model with a Python tool, effort max, 2026-08-3093.7%
  • Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0190.2%
  • GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0989.1%
  • Claude Fable 5 Model with a Python tool, effort max, 2026-06-0987.0%
  • GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0986.0%
  • Claude Opus 5 Model with a Python tool, effort max, 2026-07-2485.6%
  • GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1185.3%
  • GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0982.1%
  • Claude Opus 4.8 Model with a Python tool, effort max, 2026-06-1080.0%
  • GPT-5.4 Model with a Python tool, effort xhigh, 2026-06-1178.6%
  • Gemini 3.7 Flash Model with a Python tool, effort high, 2026-08-1471.6%
  • Claude Opus 4.7 Model with a Python tool, effort max, 2026-06-1070.2%
  • Gemini 3.8 Flash Model with a Python tool, effort high, 2026-09-0268.4%
  • GPT-5.2 Model with a Python tool, effort xhigh, 2026-06-1167.4%
  • Claude Opus 4.6 Model with a Python tool, effort max, 2026-06-1166.0%
  • Claude Sonnet 5 Model with a Python tool, effort max, 2026-06-3065.6%
  • Gemini 3.5 Flash Model with a Python tool, effort high, 2026-06-1062.8%
  • Gemini 3.1 Pro (preview) Model with a Python tool, effort default, 2026-06-1159.6%
  • Gemini 3.6 Flash Model with a Python tool, effort high, 2026-08-0258.9%
  • Gemini 3 Flash (preview) Model with a Python tool, effort default, 2026-06-1151.2%
  • GPT-5.4 Mini Model with a Python tool, effort xhigh, 2026-06-1251.2%
  • GPT-5 mini Model with a Python tool, effort high, 2026-06-1246.7%
  • GPT-5.4 Nano Model with a Python tool, effort high, 2026-06-1244.9%
  • GPT-5.6 Luna Model with a Python tool, effort low, 2026-08-2941.4%
  • GPT-5.6 Luna Model with a Python tool, effort none, 2026-08-2939.6%
  • Claude Opus 4.5 Model with a Python tool, thinking budget 32K, 2026-06-1134.4%
  • Gemini 3.5 Flash-Lite Model with a Python tool, effort high, 2026-08-0226.0%
  • GPT-5.4 Mini Model with a Python tool, effort low, 2026-08-2824.6%
  • Claude Sonnet 4.5 Model with a Python tool, thinking budget 32K, 2026-06-1123.9%
  • GPT-5.4 Nano Model with a Python tool, effort low, 2026-08-2820.4%
  • GPT-5 mini Model with a Python tool, effort low, 2026-08-2718.2%
  • GPT-5.4 Mini Model with a Python tool, effort none, 2026-08-2817.2%
  • GPT-5 mini Model with a Python tool, effort minimal, 2026-08-276.0%
  • GPT-5.4 Nano Model with a Python tool, effort none, 2026-08-284.6%

Watch out

Epoch released version 2 on 2026-06-12 after fixing errors in 42% of FrontierMath problems. Scores here are on version 2. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability not measured by this source.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute. Epoch AI says FrontierMath was developed with funding from OpenAI, which has exclusive access to a subset of the benchmark.

See Epoch AI's results
Not tested here: 8 models

Claude Haiku 4.5, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-5.1, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

More about this test: FrontierMath (tiers 1 to 3)

What's in the test

295 problems covering most major branches of modern mathematics. A typical problem takes a researcher in that field several hours. The model submits a Python function that returns its answer, which is checked automatically. Scores are on Epoch's private problems.

295 problems

Technical details for FrontierMath (tiers 1 to 3)

Accuracy: The share of the 295 problems the model solved, in Epoch AI's own runs.

  • GPT-6 Astra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Fable 5.1, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.6 Sol, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Fable 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.6 Terra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Opus 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.5, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • GPT-5.6 Luna, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Opus 4.8, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.4, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • Gemini 3.7 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Claude Opus 4.7, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Gemini 3.8 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.2, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • Claude Opus 4.6, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Sonnet 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Gemini 3.5 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Gemini 3.1 Pro (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
  • Gemini 3.6 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Gemini 3 Flash (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
  • GPT-5.4 Mini, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • GPT-5 mini, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.4 Nano, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.6 Luna, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
  • GPT-5.6 Luna, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
  • Claude Opus 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
  • Gemini 3.5 Flash-Lite, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.4 Mini, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
  • Claude Sonnet 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
  • GPT-5.4 Nano, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
  • GPT-5 mini, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
  • GPT-5.4 Mini, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
  • GPT-5 mini, model with a Python tool, effort minimal: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: minimal.
  • GPT-5.4 Nano, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2026-06-09 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

FrontierMath (tier 4)

best 97.6%Epoch AI43 problemsNewest result 2026-09-02reliability not measuredcaveat

The hardest tier of FrontierMath: exceptionally difficult research-level math problems written and checked by expert mathematicians. Epoch AI runs the test itself, and the model can write and run Python code while it works.

Results

Each model's best setting in this test

  • GPT-6 Astra Model with a Python tool, effort high, 2026-08-3097.6%
  • Claude Fable 5 Model with a Python tool, effort max, 2026-06-0990.2%
  • Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0187.8%
  • GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0982.9%
  • Claude Opus 5 Model with a Python tool, effort max, 2026-07-2473.2%
  • GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1172.5%
  • GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0970.7%
  • GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0961.0%
Show all 31 results (23 not shown above) from FrontierMath (tier 4)
  • GPT-6 Astra Model with a Python tool, effort high, 2026-08-3097.6%
  • GPT-6 Astra Model with a Python tool, effort max, 2026-08-3097.6%
  • GPT-6 Astra Model with a Python tool, effort xhigh, 2026-08-3097.6%
  • GPT-6 Astra Model with a Python tool, effort medium, 2026-08-3097.6%
  • Claude Fable 5 Model with a Python tool, effort max, 2026-06-0990.2%
  • Claude Fable 5.1 Model with a Python tool, effort max, 2026-09-0187.8%
  • GPT-6 Astra Model with a Python tool, effort low, 2026-08-3087.8%
  • GPT-5.6 Sol Model with a Python tool, effort max, 2026-07-0982.9%
  • GPT-6 Astra Model with a Python tool, effort none, 2026-08-3082.9%
  • Claude Opus 5 Model with a Python tool, effort max, 2026-07-2473.2%
  • GPT-5.5 Model with a Python tool, effort xhigh, 2026-06-1172.5%
  • GPT-5.6 Terra Model with a Python tool, effort max, 2026-07-0970.7%
  • GPT-5.6 Luna Model with a Python tool, effort max, 2026-07-0961.0%
  • Claude Opus 4.8 Model with a Python tool, effort max, 2026-06-1056.1%
  • GPT-5.4 Model with a Python tool, effort xhigh, 2026-06-1149.0%
  • Gemini 3.7 Flash Model with a Python tool, effort high, 2026-08-1436.6%
  • Claude Opus 4.7 Model with a Python tool, effort max, 2026-06-1031.7%
  • GPT-5.2 Model with a Python tool, effort xhigh, 2026-06-1131.7%
  • Claude Sonnet 5 Model with a Python tool, effort max, 2026-06-3029.3%
  • Claude Opus 4.6 Model with a Python tool, effort max, 2026-06-1126.8%
  • Gemini 3.1 Pro (preview) Model with a Python tool, effort default, 2026-06-1126.8%
  • Gemini 3.5 Flash Model with a Python tool, effort high, 2026-06-1026.8%
  • Gemini 3.6 Flash Model with a Python tool, effort high, 2026-08-0222.0%
  • Gemini 3.8 Flash Model with a Python tool, effort high, 2026-09-0222.0%
  • Gemini 3 Flash (preview) Model with a Python tool, effort default, 2026-06-1117.1%
  • GPT-5 mini Model with a Python tool, effort high, 2026-06-1212.2%
  • GPT-5.4 Nano Model with a Python tool, effort high, 2026-06-1212.2%
  • GPT-5.4 Mini Model with a Python tool, effort xhigh, 2026-06-129.8%
  • Claude Opus 4.5 Model with a Python tool, thinking budget 32K, 2026-06-114.9%
  • Claude Sonnet 4.5 Model with a Python tool, thinking budget 32K, 2026-06-112.4%
  • Gemini 3.5 Flash-Lite Model with a Python tool, effort high, 2026-08-020.0%

Watch out

With only 43 problems, each one is worth more than 2 points. Epoch released version 2 on 2026-06-12 after fixing errors in 42% of FrontierMath problems. Scores here are on version 2. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability not measured by this source.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute. Epoch AI says FrontierMath was developed with funding from OpenAI, which has exclusive access to a subset of the benchmark.

See Epoch AI's results
Not tested here: 8 models

Claude Haiku 4.5, Claude Opus 5.5, Claude Sonnet 4.6, Gemini 3 Pro (preview), GPT-5.1, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

More about this test: FrontierMath (tier 4)

What's in the test

43 problems. The hardest can take a researcher in that field several days. The model submits a Python function that returns its answer, which is checked automatically. Scores are on Epoch's private problems.

43 problems

Technical details for FrontierMath (tier 4)

Accuracy: The share of the 43 problems the model solved, in Epoch AI's own runs.

  • GPT-6 Astra, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-6 Astra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-6 Astra, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • GPT-6 Astra, model with a Python tool, effort medium: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: medium.
  • Claude Fable 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Fable 5.1, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-6 Astra, model with a Python tool, effort low: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: low.
  • GPT-5.6 Sol, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-6 Astra, model with a Python tool, effort none: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: none.
  • Claude Opus 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.5, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • GPT-5.6 Terra, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.6 Luna, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Opus 4.8, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.4, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • Gemini 3.7 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Claude Opus 4.7, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • GPT-5.2, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • Claude Sonnet 5, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Claude Opus 4.6, model with a Python tool, effort max: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: max.
  • Gemini 3.1 Pro (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
  • Gemini 3.5 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Gemini 3.6 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Gemini 3.8 Flash, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • Gemini 3 Flash (preview), model with a Python tool, effort default: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: default.
  • GPT-5 mini, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.4 Nano, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.
  • GPT-5.4 Mini, model with a Python tool, effort xhigh: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: xhigh.
  • Claude Opus 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
  • Claude Sonnet 4.5, model with a Python tool, thinking budget 32K: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Thinking budget: 32K tokens.
  • Gemini 3.5 Flash-Lite, model with a Python tool, effort high: Epoch AI's evaluation harness (Inspect), with a Python tool. Tools: A Python tool for running code; the answer is submitted as a Python function. Input: Epoch AI's copy of the test problems. Effort: high.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2026-06-09 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

OTIS Mock AIME

best 100.0%Epoch AI45 problemsNewest result 2026-09-02repeats runs, method under the sourcecaveat

Competition-style math problems from the OTIS Mock AIME exams, written by students in the OTIS olympiad training program. Epoch AI runs the test itself.

Results

Each model's best setting in this test

  • Claude Fable 5 Sent straight to the model, effort high, 2026-08-06100.0%
  • Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-01100.0%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-09100.0%
  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-30100.0%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0999.7%
  • Claude Opus 5 Sent straight to the model, effort max, 2026-07-2498.9%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0298.9%
  • Claude Opus 4.8 Sent straight to the model, effort max, 2026-06-0798.3%
  • GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0998.3%
Show all 80 results (71 not shown above) from OTIS Mock AIME
  • Claude Fable 5 Sent straight to the model, effort high, 2026-08-06100.0%
  • Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-01100.0%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-09100.0%
  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-30100.0%
  • Claude Fable 5 Sent straight to the model, effort max, 2026-06-1099.7%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0999.7%
  • Claude Opus 5 Sent straight to the model, effort max, 2026-07-2498.9%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0298.9%
  • Claude Opus 4.8 Sent straight to the model, effort max, 2026-06-0798.3%
  • GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0998.3%
  • Claude Opus 4.7 Sent straight to the model, effort xhigh, 2026-04-1797.8%
  • Claude Fable 5 Sent straight to the model, effort low, 2026-08-0697.8%
  • Claude Opus 4.8 Sent straight to the model, effort low, 2026-08-0697.8%
  • Claude Opus 5 Sent straight to the model, effort default, 2026-08-0697.8%
  • GPT-5.4 Sent straight to the model, effort high, 2026-07-1597.8%
  • Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1497.2%
  • GPT-5.2 Sent straight to the model, effort high, 2025-12-1196.1%
  • GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1396.1%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-2095.6%
  • Gemini 3 Flash (preview) Sent straight to the model, effort high, 2026-08-0695.6%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0695.6%
  • Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2595.6%
  • GPT-5.4 Sent straight to the model, effort medium, 2026-07-1595.6%
  • GPT-5.6 Sol Sent straight to the model, effort low, 2026-08-0795.6%
  • GPT-5.4 Sent straight to the model, effort xhigh, 2026-03-0695.3%
  • Claude Sonnet 5 Sent straight to the model, effort xhigh, 2026-07-0194.7%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 64K, 2026-02-0694.4%
  • Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0294.2%
  • GPT-5.2 Sent straight to the model, effort medium, 2025-12-1193.9%
  • Claude Opus 5 Sent straight to the model, effort low, 2026-08-0693.3%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 32K, 2026-02-0693.1%
  • Gemini 3 Flash (preview) Sent straight to the model, effort default, 2025-12-1792.8%
  • Gemini 3 Pro (preview) Sent straight to the model, effort default, 2025-11-1991.4%
  • Claude Opus 4.6 Sent straight to the model, effort max, 2026-08-0691.1%
  • Gemini 3.5 Flash Sent straight to the model, effort low, 2026-08-0688.9%
  • GPT-5.4 Mini Sent straight to the model, effort xhigh, 2026-08-0788.9%
  • GPT-5.6 Terra Sent straight to the model, effort low, 2026-08-0788.9%
  • GPT-5.1 Sent straight to the model, effort high, 2025-11-1388.6%
  • GPT-5.4 Nano Sent straight to the model, effort high, 2026-04-1487.8%
  • GPT-5.4 Mini Sent straight to the model, effort high, 2026-04-1587.2%
  • Claude Opus 4.7 Sent straight to the model, effort max, 2026-08-0686.7%
  • GPT-5 mini Sent straight to the model, effort high, 2025-10-3086.7%
  • Claude Opus 4.5 Sent straight to the model, thinking budget 32K, 2025-11-2486.1%
  • Claude Sonnet 4.6 Sent straight to the model, thinking budget 32K, 2026-02-2085.8%
  • GPT-5.1 Sent straight to the model, effort medium, 2025-11-1785.6%
  • Claude Opus 4.8 Sent straight to the model, effort none, 2026-08-0684.4%
  • GPT-5.4 Sent straight to the model, effort low, 2026-07-1584.4%
  • GPT-5.5 Sent straight to the model, effort low, 2026-08-0784.4%
  • Claude Sonnet 4.6 Sent straight to the model, effort medium, 2026-07-1382.2%
  • Gemini 3.6 Flash Sent straight to the model, effort low, 2026-08-0782.2%
  • Claude Opus 4.5 Sent straight to the model, thinking budget 16K, 2025-11-2481.7%
  • Claude Sonnet 5 Sent straight to the model, effort max, 2026-08-0680.0%
  • Gemini 3.5 Flash Sent straight to the model, effort minimal, 2026-07-1580.0%
  • Gemini 3.6 Flash Sent straight to the model, effort minimal, 2026-08-0780.0%
  • GPT-5.2 Sent straight to the model, effort low, 2025-12-1178.9%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2177.8%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 59K, 2025-10-2877.8%
  • Claude Sonnet 4.6 Sent straight to the model, effort high, 2026-07-1375.6%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 16K, 2025-10-2871.1%
  • Claude Sonnet 4.6 Sent straight to the model, effort max, 2026-08-0671.1%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort high, 2026-08-0671.1%
  • GPT-5.4 Nano Sent straight to the model, effort low, 2026-08-0768.9%
  • GPT-5.6 Sol Sent straight to the model, effort none, 2026-08-0768.9%
  • Claude Haiku 4.5 Sent straight to the model, thinking budget 32K, 2025-10-2266.7%
  • GPT-5.6 Luna Sent straight to the model, effort low, 2026-08-0766.7%
  • GPT-5.1 Sent straight to the model, effort low, 2025-11-2563.9%
  • GPT-5.2 Sent straight to the model, effort none, 2026-07-1362.2%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort low, 2026-08-0660.0%
  • GPT-5.4 Sent straight to the model, effort none, 2026-07-1557.8%
  • GPT-5.5 Sent straight to the model, effort none, 2026-08-0757.8%
  • GPT-5 mini Sent straight to the model, effort minimal, 2026-08-0755.6%
  • GPT-5.6 Terra Sent straight to the model, effort none, 2026-08-0753.3%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort minimal, 2026-08-0651.1%
  • Claude Opus 4.5 Sent straight to the model, effort default, 2025-11-2448.1%
  • GPT-5.4 Nano Sent straight to the model, effort none, 2026-08-0746.7%
  • GPT-5.6 Luna Sent straight to the model, effort none, 2026-08-0740.0%
  • GPT-5.1 Sent straight to the model, effort none, 2026-08-0737.8%
  • Claude Haiku 4.5 Sent straight to the model, effort default, 2025-10-1635.8%
  • Claude Sonnet 4.5 Sent straight to the model, effort default, 2025-09-2935.6%
  • GPT-5.4 Mini Sent straight to the model, effort none, 2026-08-0726.7%

Watch out

With 45 problems, each one is worth more than 2 points. Three problems include a picture; Epoch leaves the picture out and gives every model the code that draws it instead. In September 2025 Epoch changed how it reads the final answer, after its earlier grader marked some blank answers correct. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability: Epoch AI says it runs most models 16 times on this test and reports the average score. How many runs each model got is not in its download, and a few scores do not fit 16 runs.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.

See Epoch AI's results
Not tested here: 4 models

Claude Opus 5.5, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

More about this test: OTIS Mock AIME

What's in the test

45 problems from three mock exams (2024, 2025 I and 2025 II), 15 from each. Every answer is a whole number from 0 to 999. Epoch rates them harder than MATH Level 5 but easier than FrontierMath.

45 problems

Older results, more than a year old (1) from OTIS Mock AIME
  • GPT-5 mini Sent straight to the model, effort medium, 2025-08-0778.3%
Technical details for OTIS Mock AIME

Accuracy: The share of the 45 problems the model answered correctly, in Epoch AI's own runs.

  • Claude Fable 5, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Fable 5.1, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Sol, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-6 Astra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Fable 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Terra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Opus 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.8 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 4.8, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Luna, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Opus 4.7, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Fable 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.8, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.4, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.7 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Gemini 3 Flash (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.5 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • GPT-5.6 Sol, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Sonnet 5, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Opus 4.6, sent straight to the model, thinking budget 64K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 64K tokens.
  • Gemini 3.6 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Claude Opus 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Gemini 3 Flash (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Gemini 3 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Opus 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.5 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4 Mini, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • GPT-5.6 Terra, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.1, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4 Nano, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4 Mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 4.7, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5 mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Sonnet 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5.1, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Claude Opus 4.8, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.4, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 4.6, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Gemini 3.6 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
  • Claude Sonnet 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.5 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Gemini 3.6 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • GPT-5.2, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 59K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 59K tokens.
  • Claude Sonnet 4.6, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 16K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 16K tokens.
  • Claude Sonnet 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4 Nano, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.6 Sol, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Haiku 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5.6 Luna, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.1, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.2, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.5, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5 mini, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • GPT-5.6 Terra, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Claude Opus 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.4 Nano, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.6 Luna, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.1, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Haiku 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Sonnet 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.4 Mini, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2025-09-29 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

MathArena ArXivMath (June 2026)

best 85.4%MathArena (ETH Zurich SRI Lab and INSAIT)48 problemsNewest result 2026-09-07reliability not measuredcaveat

MathArena, from ETH Zurich, tests AI on new maths problems soon after they appear, so models are less likely to have seen them. ArXivMath takes its problems from research papers posted on arXiv each month; each problem has one correct final answer.

Results

Run date not published; posted 2026-09-07

  • Claude Fable 5 Sent straight to the model, effort max85.4%
  • GPT-5.5 Sent straight to the model, effort xhigh83.6%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort not stated66.7%
  • Gemini 3.5 Flash Sent straight to the model, effort not stated52.1%

Watch out

MathArena does not publish when it ran each model. The date shown is when MathArena published that model's answers. The number of runs differs by model and is shown with each result; MathArena's site says it runs each model 4 times, but its published answers show 3 to 7 runs per problem. MathArena marks 1 of the models shown as released after these problems were published, so it may have seen the source papers in training: Claude-Fable-5 (max). Cost is MathArena's estimate for one run on one problem at the model maker's prices. With 48 problems, each one is worth about 2 points. Results MathArena has not published answers for, or whose published answers do not match its site, are not shown. Nor are results whose settings MathArena published only after their answers, since the settings used for those runs cannot be confirmed. Data: MathArena, ETH Zurich, matharena.ai, CC BY-SA 4.0, shown with MathArena's permission.

Reliability not measured by this source.

MathArena is run by the SRI Lab at ETH Zurich and INSAIT, which lists Google and DeepMind among its supporters. Google makes Gemini, one of the model families shown. MathArena does not make any of the models it tests.

See MathArena (ETH Zurich SRI Lab and INSAIT)'s results
Not tested here: 30 models

Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-5 mini, GPT-5.1, GPT-5.2, GPT-5.2 Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.

More about this test: MathArena ArXivMath (June 2026)

What's in the test

48 research-level problems taken from maths papers submitted to arXiv in June 2026. The model must give the final answer, which is checked against the answer from the paper. MathArena publishes a new set each month; only the June 2026 set is shown, because each month is a different test.

48 problems

  • Research-level maths
  • Final-answer problems
Technical details for MathArena ArXivMath (June 2026)

Accuracy: The share of the 48 research-level maths problems the model answered correctly, averaged over MathArena's runs of each problem.

  • Claude Fable 5, sent straight to the model, effort max: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort: max.
  • GPT-5.5, sent straight to the model, effort xhigh: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort: xhigh.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort not stated: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort not stated.
  • Gemini 3.5 Flash, sent straight to the model, effort not stated: MathArena's own pipeline. Tools: None. Input: The problem as text. Effort not stated.

Source version: MathArena ArXivMath June 2026 answers at commit f782eef, downloaded 2026-09-30

Results posted 2026-09-07

Run dates not published by the tester; the dates are when results were posted.

License: CC BY-SA 4.0 (MathArena's ArXivMath data); display permitted by MathArena in writing, 2026-09-28. Checked 2026-09-30.

No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano, GPT-5.1 Codex.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use