Assurance

Reasoning

Drawing sound conclusions from the facts given: following a chain of logic, solving multi-step problems, and reading a situation correctly. A good result reaches the right conclusion for the right reasons.

Updated every Friday. Sources last checked 2026-09-26.

Key evidence rings earned, out of 4Direct sent straight to the modelAgent works through the task with tools, as the words after it sayFilled: best score in that test. Dashed: tested, figures not shown here.

At a glance

independent tester: An independent group tested this, earned published in the last year: We show their numbers, published in the last year, earned two or more testers: Two or more groups tested it, earned repeat runs, some or all models: Results checked by repeat runs, for some or all models, not yet

Vendor claims have not been collected for any category yet.

LB 92.7%LB 92.7%CHSS 72.0%CHSS 72.0%MGP 84.0%MGP 84.0%
ⓘ More about this category
  • Reliability not yet measured
  • 3 sources, 34 models with figures
  • Sources last checked 2026-09-28
  • Ring 4: no rerun evidence yet
  • LB = LiveBench, CHSS = Chess Puzzles, MGP = Mystery Game Puzzles

Who does well

TestBestBest scoreNextNext score
LiveBenchGPT-6 Astra Sent straight to the modelroughly tied, 4 setupsAlso roughly tied: Claude Fable 5.1 Sent straight to the model 91.7%; GPT-5.6 Sol Sent straight to the model 91.7%92.7%Claude Opus 5.5 Sent straight to the model92.2%
Chess PuzzlesGPT-6 Astra Sent straight to the model, effort max72.0%Gemini 3.8 Flash Sent straight to the model, effort high61.0%
Mystery Game PuzzlesGPT-6 Astra Minimal agent with a submission tool, effort max84.0%Claude Opus 5 Minimal agent with a submission tool, effort max59.0%
Place within each test. Places are never added up across tests.
ModelLiveBenchChess PuzzlesMystery Game Puzzles
Claude Fable 57 of 28Sent straight to the model13 of 30Sent straight to the model, effort high6 of 28Minimal agent with a submission tool, effort max
Claude Fable 5.13 of 28Sent straight to the model8 of 30Sent straight to the model, effort max3 of 28Minimal agent with a submission tool, effort max
Claude Opus 4.525 of 28Sent straight to the model28 of 30Sent straight to the model, thinking budget 32K20 of 28Minimal agent with a submission tool, thinking budget 48K
Claude Opus 4.612 of 28Sent straight to the model26 of 30Sent straight to the model, thinking budget 32K18 of 28Minimal agent with a submission tool, effort max
Claude Opus 4.716 of 28Sent straight to the model20 of 30Sent straight to the model, effort xhigh16 of 28Minimal agent with a submission tool, effort max
Claude Opus 4.810 of 28Sent straight to the model17 of 30Sent straight to the model, effort max10 of 28Minimal agent with a submission tool, effort max
Claude Opus 55 of 28Sent straight to the model12 of 30Sent straight to the model, effort max2 of 28Minimal agent with a submission tool, effort max
Claude Sonnet 4.528 of 30Sent straight to the model, thinking budget 32K24 of 28Minimal agent with a submission tool, thinking budget 48K
Claude Sonnet 4.619 of 28Sent straight to the model27 of 30Sent straight to the model, thinking budget 32K25 of 28Minimal agent with a submission tool, effort low
Claude Sonnet 511 of 28Sent straight to the model16 of 30Sent straight to the model, effort xhigh11 of 28Minimal agent with a submission tool, effort max
Gemini 3 Flash (preview)14 of 30Sent straight to the model, effort high17 of 28Minimal agent with a submission tool, effort low
Gemini 3.1 Pro (preview)20 of 28Sent straight to the model3 of 30Sent straight to the model, effort default13 of 28Minimal agent with a submission tool, effort high
Gemini 3.5 Flash22 of 28Sent straight to the model6 of 30Sent straight to the model, effort high14 of 28Minimal agent with a submission tool, effort high
Gemini 3.5 Flash-Lite28 of 28Sent straight to the model25 of 30Sent straight to the model, effort high22 of 28Minimal agent with a submission tool, effort low
Gemini 3.6 Flash18 of 28Sent straight to the model11 of 30Sent straight to the model, effort low15 of 28Minimal agent with a submission tool, effort high
Gemini 3.7 Flash15 of 28Sent straight to the model8 of 30Sent straight to the model, effort high8 of 28Minimal agent with a submission tool, effort high
Gemini 3.8 Flash9 of 28Sent straight to the model2 of 30Sent straight to the model, effort high7 of 28Minimal agent with a submission tool, effort high
GPT-5 mini20 of 30Sent straight to the model, effort high27 of 28Minimal agent with a submission tool, effort minimal
GPT-5.118 of 30Sent straight to the model, effort high22 of 28Minimal agent with a submission tool, effort low
GPT-5.221 of 28Sent straight to the model7 of 30Sent straight to the model, effort xhigh19 of 28Minimal agent with a submission tool, effort high
GPT-5.414 of 28Sent straight to the model10 of 30Sent straight to the model, effort xhigh8 of 28Minimal agent with a submission tool, effort xhigh
GPT-5.4 Mini27 of 28Sent straight to the model24 of 30Sent straight to the model, effort xhigh26 of 28Minimal agent with a submission tool, effort none
GPT-5.4 Nano24 of 28Sent straight to the model20 of 30Sent straight to the model, effort high28 of 28Minimal agent with a submission tool, effort none
GPT-5.57 of 28Sent straight to the model23 of 30Sent straight to the model, effort low5 of 28Minimal agent with a submission tool, effort xhigh
GPT-5.6 Luna17 of 28Sent straight to the model14 of 30Sent straight to the model, effort max21 of 28Minimal agent with a submission tool, effort max
GPT-5.6 Sol4 of 28Sent straight to the model3 of 30Sent straight to the model, effort max3 of 28Minimal agent with a submission tool, effort max
GPT-5.6 Terra6 of 28Sent straight to the model5 of 30Sent straight to the model, effort max11 of 28Minimal agent with a submission tool, effort max
GPT-6 Astra1 of 28Sent straight to the model1 of 30Sent straight to the model, effort max1 of 28Minimal agent with a submission tool, effort max

The tests here put 224 pairs of models in a different order.

Show all 224 pairs placed in a different order
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high.
  • In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
  • In Chess Puzzles, Claude Fable 5, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Fable 5, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Claude Fable 5.1, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, Claude Fable 5.1, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Claude Fable 5.1, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.6 Sol, sent straight to the model; in Chess Puzzles, GPT-5.6 Sol, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In Chess Puzzles, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.1, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.1, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.1, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in Chess Puzzles, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
  • In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Sol, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Sol, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Sol, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Mystery Game Puzzles, GPT-5.2, minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
  • In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3 Flash (preview), sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3 Flash (preview), minimal agent with a submission tool, effort low, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
  • In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.4, sent straight to the model, effort xhigh.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Terra, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Terra, sent straight to the model, effort max; in Mystery Game Puzzles, GPT-5.6 Terra, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
  • In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
  • In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort xhigh.
  • In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In LiveBench, GPT-5.4 Mini, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.6 Flash, sent straight to the model, effort low; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, effort max.
  • In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, effort max; in Mystery Game Puzzles, GPT-5.6 Sol, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Terra, sent straight to the model, effort max.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.8 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.4 Mini, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
  • In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.1, minimal agent with a submission tool, effort low.
  • In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort xhigh.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.2, minimal agent with a submission tool, effort high.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.6 Luna, sent straight to the model, effort max.
  • In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Mystery Game Puzzles, GPT-5.2, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.4 Mini, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
  • In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
  • In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
  • In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above GPT-5.5, sent straight to the model; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
  • In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.

Findings for this category aren't written yet.

The tests behind this

LiveBench

best 92.7%LiveBenchNewest result 2026-06-25reliability not measuredcaveat

A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.

Results

All Sent straight to the model · Run date not published; posted 2026-06-25

Each model's best setting in this test

  • GPT-6 Astra92.7%
  • Claude Opus 5.592.2%
  • Claude Fable 5.191.7%
  • GPT-5.6 Sol91.7%
  • Claude Opus 591.2%
  • GPT-5.6 Terra90.6%
  • Claude Fable 589.7%
  • GPT-5.589.7%
Show all 29 results (21 not shown above) from LiveBench
  • GPT-6 Astra92.7%
  • Claude Opus 5.592.2%
  • Claude Fable 5.191.7%
  • GPT-5.6 Sol91.7%
  • Claude Opus 591.2%
  • Claude Opus 5.590.7%
  • GPT-5.6 Terra90.6%
  • Claude Fable 589.7%
  • GPT-5.589.7%
  • Gemini 3.8 Flash89.3%
  • Claude Opus 4.889.2%
  • Claude Sonnet 588.7%
  • Claude Opus 4.688.7%
  • GPT-6 Sol88.7%
  • GPT-5.488.1%
  • Gemini 3.7 Flash87.8%
  • Claude Opus 4.787.2%
  • GPT-5.6 Luna85.6%
  • Gemini 3.6 Flash85.2%
  • Claude Sonnet 4.684.8%
  • Gemini 3.1 Pro (preview)84.0%
  • GPT-5.283.2%
  • Gemini 3.5 Flash82.0%
  • GPT-6 Luna81.8%
  • GPT-5.4 Nano81.1%
  • Claude Opus 4.580.1%
  • GPT-5.2 Codex77.7%
  • GPT-5.4 Mini71.3%
  • Gemini 3.5 Flash-Lite60.2%

Watch out

LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.

Reliability not measured by this source.

LiveBench is funded by Abacus.AI.

See LiveBench's results
Not tested here: 6 models

Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1.

More about this test: LiveBench

What's in the test

Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.

Technical details for LiveBench

Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.

  • GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
  • GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
  • Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
  • GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
  • GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
  • GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
  • GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
  • Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.

Source version: LiveBench release 2026-06-25

Results posted 2026-06-25

Run dates not published by the tester; the dates are when results were posted.

License: Apache License 2.0. Checked 2026-09-27.

Chess Puzzles

best 72.0%Epoch AI100 puzzlesNewest result 2026-09-02reliability not measuredcaveat

Chess positions where the model must find the single best next move, as judged by the Stockfish chess engine. Epoch AI made the puzzles and runs the test itself.

Results

Each model's best setting in this test

  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-3072.0%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0261.0%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-1955.0%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0955.0%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0954.0%
  • Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2850.0%
  • GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1549.0%
  • Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-0147.0%
  • Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1447.0%
Show all 78 results (69 not shown above) from Chess Puzzles
  • GPT-6 Astra Sent straight to the model, effort max, 2026-08-3072.0%
  • Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0261.0%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-1955.0%
  • GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0955.0%
  • GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0954.0%
  • Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2850.0%
  • Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0649.0%
  • GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1549.0%
  • Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-0147.0%
  • Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1447.0%
  • Gemini 3.5 Flash Sent straight to the model, effort low, 2026-08-0645.0%
  • GPT-5.4 Sent straight to the model, effort xhigh, 2026-03-1144.0%
  • Gemini 3.5 Flash Sent straight to the model, effort minimal, 2026-07-1543.0%
  • Gemini 3.6 Flash Sent straight to the model, effort low, 2026-08-0743.0%
  • Claude Opus 5 Sent straight to the model, effort max, 2026-07-2442.0%
  • Claude Fable 5 Sent straight to the model, effort high, 2026-08-0641.0%
  • Claude Fable 5 Sent straight to the model, effort max, 2026-06-0941.0%
  • Gemini 3 Flash (preview) Sent straight to the model, effort high, 2026-08-0640.0%
  • Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0240.0%
  • GPT-5.2 Sent straight to the model, effort high, 2025-12-1140.0%
  • GPT-5.2 Sent straight to the model, effort medium, 2025-12-1140.0%
  • GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0940.0%
  • Gemini 3 Flash (preview) Sent straight to the model, effort default, 2025-12-1738.0%
  • GPT-5.4 Sent straight to the model, effort high, 2026-07-1538.0%
  • GPT-5.4 Sent straight to the model, effort medium, 2026-07-1538.0%
  • Claude Sonnet 5 Sent straight to the model, effort xhigh, 2026-06-3035.0%
  • Gemini 3.6 Flash Sent straight to the model, effort minimal, 2026-08-0735.0%
  • Claude Opus 4.8 Sent straight to the model, effort max, 2026-05-2934.0%
  • Claude Opus 5 Sent straight to the model, effort default, 2026-08-0633.0%
  • GPT-5.1 Sent straight to the model, effort high, 2025-12-0832.0%
  • Gemini 3 Pro (preview) Sent straight to the model, effort default, 2025-12-0831.0%
  • Claude Opus 4.7 Sent straight to the model, effort xhigh, 2026-04-2030.0%
  • GPT-5 mini Sent straight to the model, effort high, 2026-08-0730.0%
  • GPT-5.4 Nano Sent straight to the model, effort high, 2026-04-1430.0%
  • Claude Opus 4.8 Sent straight to the model, effort low, 2026-08-0629.0%
  • Claude Fable 5 Sent straight to the model, effort low, 2026-08-0628.0%
  • GPT-5.6 Sol Sent straight to the model, effort low, 2026-08-0727.0%
  • GPT-5.5 Sent straight to the model, effort low, 2026-08-0726.0%
  • GPT-5.4 Mini Sent straight to the model, effort xhigh, 2026-08-0724.0%
  • GPT-5.2 Sent straight to the model, effort low, 2025-12-1123.0%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort high, 2026-08-0622.0%
  • GPT-5.6 Terra Sent straight to the model, effort low, 2026-08-0722.0%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort minimal, 2026-08-0621.0%
  • GPT-5.6 Luna Sent straight to the model, effort low, 2026-08-0721.0%
  • Claude Opus 4.7 Sent straight to the model, effort low, 2026-07-1420.0%
  • Claude Opus 5 Sent straight to the model, effort low, 2026-08-0620.0%
  • GPT-5.4 Sent straight to the model, effort low, 2026-07-1520.0%
  • Gemini 3.5 Flash-Lite Sent straight to the model, effort low, 2026-08-0618.0%
  • GPT-5.4 Mini Sent straight to the model, effort high, 2026-04-1518.0%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 32K, 2026-02-0617.0%
  • GPT-5.1 Sent straight to the model, effort none, 2026-08-0717.0%
  • GPT-5.4 Nano Sent straight to the model, effort low, 2026-08-0717.0%
  • Claude Sonnet 5 Sent straight to the model, effort max, 2026-06-3016.0%
  • Claude Opus 4.6 Sent straight to the model, effort max, 2026-08-0614.0%
  • GPT-5.1 Sent straight to the model, effort low, 2026-08-0714.0%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 120K, 2026-02-2013.0%
  • Claude Opus 4.8 Sent straight to the model, effort none, 2026-08-0613.0%
  • Claude Sonnet 4.6 Sent straight to the model, thinking budget 32K, 2026-02-2013.0%
  • Claude Opus 4.5 Sent straight to the model, thinking budget 32K, 2025-12-0812.0%
  • Claude Sonnet 4.5 Sent straight to the model, thinking budget 32K, 2025-12-0812.0%
  • GPT-5 mini Sent straight to the model, effort low, 2026-08-0712.0%
  • Claude Opus 4.6 Sent straight to the model, thinking budget 64K, 2026-02-0610.0%
  • GPT-5.5 Sent straight to the model, effort none, 2026-08-0710.0%
  • Claude Haiku 4.5 Sent straight to the model, thinking budget 32K, 2026-07-168.0%
  • Claude Sonnet 4.6 Sent straight to the model, effort medium, 2026-07-138.0%
  • Claude Opus 4.7 Sent straight to the model, effort max, 2026-08-067.0%
  • GPT-5 mini Sent straight to the model, effort minimal, 2026-08-077.0%
  • GPT-5.6 Sol Sent straight to the model, effort none, 2026-08-077.0%
  • Claude Sonnet 4.6 Sent straight to the model, effort high, 2026-07-135.0%
  • GPT-5.4 Sent straight to the model, effort none, 2026-07-155.0%
  • GPT-5.6 Terra Sent straight to the model, effort none, 2026-08-075.0%
  • Claude Opus 4.5 Sent straight to the model, effort default, 2026-08-064.0%
  • Claude Sonnet 4.5 Sent straight to the model, effort default, 2026-08-064.0%
  • GPT-5.2 Sent straight to the model, effort none, 2026-07-134.0%
  • Claude Sonnet 4.6 Sent straight to the model, effort max, 2026-08-063.0%
  • GPT-5.4 Mini Sent straight to the model, effort none, 2026-08-073.0%
  • GPT-5.4 Nano Sent straight to the model, effort none, 2026-08-073.0%
  • GPT-5.6 Luna Sent straight to the model, effort none, 2026-08-072.0%

Watch out

Epoch says chess itself matters little; it uses these puzzles as a rough measure of spatial reasoning and planning. When Epoch tested five human players on 20 puzzles each, the two best got 75% right, more than the best model on the same puzzles. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability not measured by this source.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.

See Epoch AI's results
Not tested here: 4 models

Claude Opus 5.5, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

More about this test: Chess Puzzles

What's in the test

100 new puzzles made by a program, so they do not appear anywhere else. The model gets the board as text in standard chess notation, not as a picture, and answers with one move.

100 puzzles

Technical details for Chess Puzzles

Accuracy: The share of the 100 puzzles where the model chose the best move, in Epoch AI's own runs.

  • GPT-6 Astra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.8 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.6 Sol, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.6 Terra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.5 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.1 Pro (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Claude Fable 5.1, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3.7 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.5 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Gemini 3.5 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Gemini 3.6 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Fable 5, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Fable 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3 Flash (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3.6 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.2, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • GPT-5.6 Luna, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Gemini 3 Flash (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.4, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Claude Sonnet 5, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • Gemini 3.6 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • Claude Opus 4.8, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Opus 5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.1, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Gemini 3 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Opus 4.7, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • GPT-5 mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4 Nano, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 4.8, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Fable 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.6 Sol, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4 Mini, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
  • GPT-5.2, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.6 Terra, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • GPT-5.6 Luna, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.7, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Gemini 3.5 Flash-Lite, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • GPT-5.4 Mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • Claude Opus 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5.1, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.4 Nano, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Sonnet 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • Claude Opus 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.1, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.6, sent straight to the model, thinking budget 120K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 120K tokens.
  • Claude Opus 4.8, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Sonnet 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Opus 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Sonnet 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • GPT-5 mini, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
  • Claude Opus 4.6, sent straight to the model, thinking budget 64K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 64K tokens.
  • GPT-5.5, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Haiku 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
  • Claude Sonnet 4.6, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
  • Claude Opus 4.7, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5 mini, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
  • GPT-5.6 Sol, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Sonnet 4.6, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
  • GPT-5.4, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.6 Terra, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Opus 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • Claude Sonnet 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
  • GPT-5.2, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • Claude Sonnet 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
  • GPT-5.4 Mini, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.4 Nano, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
  • GPT-5.6 Luna, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2025-12-08 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

Mystery Game Puzzles

best 84.0%Epoch AI100 puzzlesNewest result 2026-09-02reliability not measuredcaveat

Positions from a puzzle-oriented variant of a well-known game, where the model must pick the single best next move. Epoch AI keeps the game's name secret so that no one can prepare for the test.

Results

Each model's best setting in this test

  • GPT-6 Astra Minimal agent with a submission tool, effort max, 2026-08-3084.0%
  • Claude Opus 5 Minimal agent with a submission tool, effort max, 2026-07-2559.0%
  • Claude Fable 5.1 Minimal agent with a submission tool, effort max, 2026-09-0158.0%
  • GPT-5.6 Sol Minimal agent with a submission tool, effort max, 2026-07-2858.0%
  • GPT-5.5 Minimal agent with a submission tool, effort xhigh, 2026-07-2456.0%
  • Claude Fable 5 Minimal agent with a submission tool, effort max, 2026-07-3152.0%
  • Gemini 3.8 Flash Minimal agent with a submission tool, effort high, 2026-09-0247.0%
  • Gemini 3.7 Flash Minimal agent with a submission tool, effort high, 2026-08-1437.0%
  • GPT-5.4 Minimal agent with a submission tool, effort xhigh, 2026-07-2437.0%
Show all 69 results (60 not shown above) from Mystery Game Puzzles
  • GPT-6 Astra Minimal agent with a submission tool, effort max, 2026-08-3084.0%
  • Claude Opus 5 Minimal agent with a submission tool, effort max, 2026-07-2559.0%
  • Claude Fable 5.1 Minimal agent with a submission tool, effort max, 2026-09-0158.0%
  • GPT-5.6 Sol Minimal agent with a submission tool, effort max, 2026-07-2858.0%
  • GPT-5.5 Minimal agent with a submission tool, effort xhigh, 2026-07-2456.0%
  • Claude Fable 5 Minimal agent with a submission tool, effort max, 2026-07-3152.0%
  • GPT-5.5 Minimal agent with a submission tool, effort high, 2026-07-2752.0%
  • Gemini 3.8 Flash Minimal agent with a submission tool, effort high, 2026-09-0247.0%
  • Claude Opus 5 Minimal agent with a submission tool, effort default, 2026-08-0637.0%
  • Gemini 3.7 Flash Minimal agent with a submission tool, effort high, 2026-08-1437.0%
  • GPT-5.4 Minimal agent with a submission tool, effort xhigh, 2026-07-2437.0%
  • Claude Opus 4.8 Minimal agent with a submission tool, effort max, 2026-07-2536.0%
  • Claude Sonnet 5 Minimal agent with a submission tool, effort max, 2026-07-2835.0%
  • GPT-5.6 Terra Minimal agent with a submission tool, effort max, 2026-07-2835.0%
  • Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort high, 2026-07-2734.0%
  • GPT-5.6 Sol Minimal agent with a submission tool, effort none, 2026-08-2733.0%
  • Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort medium, 2026-08-0632.0%
  • Gemini 3.5 Flash Minimal agent with a submission tool, effort high, 2026-07-2732.0%
  • Claude Opus 4.8 Minimal agent with a submission tool, effort xhigh, 2026-07-2631.0%
  • Gemini 3.6 Flash Minimal agent with a submission tool, effort high, 2026-08-0530.0%
  • Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort low, 2026-08-0629.0%
  • Claude Opus 4.7 Minimal agent with a submission tool, effort max, 2026-07-2528.0%
  • Gemini 3.5 Flash Minimal agent with a submission tool, effort low, 2026-08-0528.0%
  • GPT-5.4 Minimal agent with a submission tool, effort medium, 2026-08-2828.0%
  • GPT-5.5 Minimal agent with a submission tool, effort low, 2026-08-2828.0%
  • Gemini 3 Flash (preview) Minimal agent with a submission tool, effort low, 2026-08-0526.0%
  • GPT-5.6 Sol Minimal agent with a submission tool, effort low, 2026-08-2726.0%
  • Claude Opus 4.6 Minimal agent with a submission tool, effort max, 2026-07-2525.0%
  • Gemini 3 Flash (preview) Minimal agent with a submission tool, effort minimal, 2026-08-0525.0%
  • Gemini 3.6 Flash Minimal agent with a submission tool, effort minimal, 2026-08-0525.0%
  • Gemini 3.6 Flash Minimal agent with a submission tool, effort low, 2026-08-0523.0%
  • GPT-5.2 Minimal agent with a submission tool, effort high, 2026-08-0623.0%
  • Claude Opus 4.5 Minimal agent with a submission tool, thinking budget 48K, 2026-07-2522.0%
  • GPT-5.2 Minimal agent with a submission tool, effort medium, 2026-08-2822.0%
  • GPT-5.6 Luna Minimal agent with a submission tool, effort max, 2026-07-2821.0%
  • Gemini 3 Flash (preview) Minimal agent with a submission tool, effort high, 2026-08-0520.0%
  • GPT-5.6 Luna Minimal agent with a submission tool, effort none, 2026-08-3020.0%
  • Gemini 3.5 Flash-Lite Minimal agent with a submission tool, effort low, 2026-08-0519.0%
  • GPT-5.1 Minimal agent with a submission tool, effort low, 2026-08-2719.0%
  • GPT-5.6 Terra Minimal agent with a submission tool, effort low, 2026-08-2719.0%
  • GPT-5.5 Minimal agent with a submission tool, effort none, 2026-08-2718.0%
  • Claude Sonnet 4.5 Minimal agent with a submission tool, thinking budget 48K, 2026-07-2517.0%
  • GPT-5.4 Minimal agent with a submission tool, effort low, 2026-08-2717.0%
  • GPT-5.6 Luna Minimal agent with a submission tool, effort low, 2026-08-2717.0%
  • Claude Sonnet 4.6 Minimal agent with a submission tool, effort low, 2026-08-0616.0%
  • Claude Sonnet 5 Minimal agent with a submission tool, effort default, 2026-08-0616.0%
  • GPT-5.1 Minimal agent with a submission tool, effort medium, 2026-08-2716.0%
  • GPT-5.4 Minimal agent with a submission tool, effort none, 2026-08-3016.0%
  • Claude Opus 4.6 Minimal agent with a submission tool, effort default, 2026-08-0615.0%
  • GPT-5.1 Minimal agent with a submission tool, effort none, 2026-08-2715.0%
  • Claude Sonnet 4.6 Minimal agent with a submission tool, effort default, 2026-08-0614.0%
  • GPT-5.2 Minimal agent with a submission tool, effort none, 2026-08-2714.0%
  • GPT-5.6 Terra Minimal agent with a submission tool, effort none, 2026-08-2714.0%
  • Claude Opus 4.7 Minimal agent with a submission tool, effort default, 2026-08-0613.0%
  • Gemini 3.5 Flash-Lite Minimal agent with a submission tool, effort minimal, 2026-08-0512.0%
  • GPT-5.6 Luna Minimal agent with a submission tool, effort medium, 2026-08-2712.0%
  • GPT-5.4 Mini Minimal agent with a submission tool, effort none, 2026-08-2711.0%
  • GPT-5 mini Minimal agent with a submission tool, effort minimal, 2026-08-2710.0%
  • GPT-5.2 Minimal agent with a submission tool, effort low, 2026-08-2710.0%
  • GPT-5.6 Terra Minimal agent with a submission tool, effort medium, 2026-08-2810.0%
  • GPT-5.4 Nano Minimal agent with a submission tool, effort none, 2026-08-279.0%
  • GPT-5.4 Mini Minimal agent with a submission tool, effort low, 2026-08-278.0%
  • Claude Opus 4.6 Minimal agent with a submission tool, effort low, 2026-08-067.0%
  • GPT-5.4 Mini Minimal agent with a submission tool, effort medium, 2026-08-277.0%
  • GPT-5.4 Nano Minimal agent with a submission tool, effort medium, 2026-08-276.0%
  • GPT-5 mini Minimal agent with a submission tool, effort high, 2026-08-275.0%
  • GPT-5.4 Nano Minimal agent with a submission tool, effort high, 2026-08-275.0%
  • GPT-5 mini Minimal agent with a submission tool, effort medium, 2026-08-274.0%
  • GPT-5.4 Nano Minimal agent with a submission tool, effort low, 2026-08-273.0%

Watch out

Epoch does not publish the prompt, example positions or model answers, so readers cannot inspect the questions or answers. An answer with no valid move gets no credit. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.

Reliability not measured by this source.

Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.

See Epoch AI's results
Not tested here: 6 models

Claude Haiku 4.5, Claude Opus 5.5, Gemini 3 Pro (preview), GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.

More about this test: Mystery Game Puzzles

What's in the test

100 positions made by a program from random games, kept only where there is a single best move. The model gets the position as text and submits its move with a tool.

100 puzzles

Technical details for Mystery Game Puzzles

Accuracy: The share of the 100 puzzles where the model chose the best move, in Epoch AI's own runs.

  • GPT-6 Astra, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Claude Opus 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Claude Fable 5.1, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • GPT-5.6 Sol, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • GPT-5.5, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
  • Claude Fable 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • GPT-5.5, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • Gemini 3.8 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • Claude Opus 5, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
  • Gemini 3.7 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • GPT-5.4, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
  • Claude Opus 4.8, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Claude Sonnet 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • GPT-5.6 Terra, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • GPT-5.6 Sol, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • Gemini 3.5 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • Claude Opus 4.8, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
  • Gemini 3.6 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Claude Opus 4.7, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Gemini 3.5 Flash, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.4, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.5, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Gemini 3 Flash (preview), minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.6 Sol, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Claude Opus 4.6, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Gemini 3 Flash (preview), minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
  • Gemini 3.6 Flash, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
  • Gemini 3.6 Flash, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.2, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Thinking budget: 48K tokens.
  • GPT-5.2, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.6 Luna, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
  • Gemini 3 Flash (preview), minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • GPT-5.6 Luna, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.1, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.6 Terra, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.5, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Thinking budget: 48K tokens.
  • GPT-5.4, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.6 Luna, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Claude Sonnet 4.6, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Claude Sonnet 5, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
  • GPT-5.1, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.4, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Claude Opus 4.6, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
  • GPT-5.1, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Claude Sonnet 4.6, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
  • GPT-5.2, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • GPT-5.6 Terra, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • Claude Opus 4.7, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
  • Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
  • GPT-5.6 Luna, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.4 Mini, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • GPT-5 mini, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
  • GPT-5.2, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.6 Terra, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.4 Nano, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
  • GPT-5.4 Mini, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • Claude Opus 4.6, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
  • GPT-5.4 Mini, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.4 Nano, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5 mini, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • GPT-5.4 Nano, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
  • GPT-5 mini, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
  • GPT-5.4 Nano, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.

Source version: Epoch AI benchmark data, downloaded 2026-09-28

Run 2026-07-24 to 2026-09-02

License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.

No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano, GPT-5.1 Codex.

What this doesn't tell you

Data version 2026-09-30+832fb354fcbe · Terms of use