Reasoning
Drawing sound conclusions from the facts given: following a chain of logic, solving multi-step problems, and reading a situation correctly. A good result reaches the right conclusion for the right reasons.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- Reliability not yet measured
- 3 sources, 34 models with figures
- Sources last checked 2026-09-28
- Ring 4: no rerun evidence yet
- LB = LiveBench, CHSS = Chess Puzzles, MGP = Mystery Game Puzzles
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| LiveBench | GPT-6 Astra Sent straight to the modelroughly tied, 4 setupsAlso roughly tied: Claude Fable 5.1 Sent straight to the model 91.7%; GPT-5.6 Sol Sent straight to the model 91.7% | 92.7% | Claude Opus 5.5 Sent straight to the model | 92.2% |
| Chess Puzzles | GPT-6 Astra Sent straight to the model, effort max | 72.0% | Gemini 3.8 Flash Sent straight to the model, effort high | 61.0% |
| Mystery Game Puzzles | GPT-6 Astra Minimal agent with a submission tool, effort max | 84.0% | Claude Opus 5 Minimal agent with a submission tool, effort max | 59.0% |
| Model | LiveBench | Chess Puzzles | Mystery Game Puzzles |
|---|---|---|---|
| Claude Fable 5 | 7 of 28Sent straight to the model | 13 of 30Sent straight to the model, effort high | 6 of 28Minimal agent with a submission tool, effort max |
| Claude Fable 5.1 | 3 of 28Sent straight to the model | 8 of 30Sent straight to the model, effort max | 3 of 28Minimal agent with a submission tool, effort max |
| Claude Opus 4.5 | 25 of 28Sent straight to the model | 28 of 30Sent straight to the model, thinking budget 32K | 20 of 28Minimal agent with a submission tool, thinking budget 48K |
| Claude Opus 4.6 | 12 of 28Sent straight to the model | 26 of 30Sent straight to the model, thinking budget 32K | 18 of 28Minimal agent with a submission tool, effort max |
| Claude Opus 4.7 | 16 of 28Sent straight to the model | 20 of 30Sent straight to the model, effort xhigh | 16 of 28Minimal agent with a submission tool, effort max |
| Claude Opus 4.8 | 10 of 28Sent straight to the model | 17 of 30Sent straight to the model, effort max | 10 of 28Minimal agent with a submission tool, effort max |
| Claude Opus 5 | 5 of 28Sent straight to the model | 12 of 30Sent straight to the model, effort max | 2 of 28Minimal agent with a submission tool, effort max |
| Claude Sonnet 4.5 | 28 of 30Sent straight to the model, thinking budget 32K | 24 of 28Minimal agent with a submission tool, thinking budget 48K | |
| Claude Sonnet 4.6 | 19 of 28Sent straight to the model | 27 of 30Sent straight to the model, thinking budget 32K | 25 of 28Minimal agent with a submission tool, effort low |
| Claude Sonnet 5 | 11 of 28Sent straight to the model | 16 of 30Sent straight to the model, effort xhigh | 11 of 28Minimal agent with a submission tool, effort max |
| Gemini 3 Flash (preview) | 14 of 30Sent straight to the model, effort high | 17 of 28Minimal agent with a submission tool, effort low | |
| Gemini 3.1 Pro (preview) | 20 of 28Sent straight to the model | 3 of 30Sent straight to the model, effort default | 13 of 28Minimal agent with a submission tool, effort high |
| Gemini 3.5 Flash | 22 of 28Sent straight to the model | 6 of 30Sent straight to the model, effort high | 14 of 28Minimal agent with a submission tool, effort high |
| Gemini 3.5 Flash-Lite | 28 of 28Sent straight to the model | 25 of 30Sent straight to the model, effort high | 22 of 28Minimal agent with a submission tool, effort low |
| Gemini 3.6 Flash | 18 of 28Sent straight to the model | 11 of 30Sent straight to the model, effort low | 15 of 28Minimal agent with a submission tool, effort high |
| Gemini 3.7 Flash | 15 of 28Sent straight to the model | 8 of 30Sent straight to the model, effort high | 8 of 28Minimal agent with a submission tool, effort high |
| Gemini 3.8 Flash | 9 of 28Sent straight to the model | 2 of 30Sent straight to the model, effort high | 7 of 28Minimal agent with a submission tool, effort high |
| GPT-5 mini | 20 of 30Sent straight to the model, effort high | 27 of 28Minimal agent with a submission tool, effort minimal | |
| GPT-5.1 | 18 of 30Sent straight to the model, effort high | 22 of 28Minimal agent with a submission tool, effort low | |
| GPT-5.2 | 21 of 28Sent straight to the model | 7 of 30Sent straight to the model, effort xhigh | 19 of 28Minimal agent with a submission tool, effort high |
| GPT-5.4 | 14 of 28Sent straight to the model | 10 of 30Sent straight to the model, effort xhigh | 8 of 28Minimal agent with a submission tool, effort xhigh |
| GPT-5.4 Mini | 27 of 28Sent straight to the model | 24 of 30Sent straight to the model, effort xhigh | 26 of 28Minimal agent with a submission tool, effort none |
| GPT-5.4 Nano | 24 of 28Sent straight to the model | 20 of 30Sent straight to the model, effort high | 28 of 28Minimal agent with a submission tool, effort none |
| GPT-5.5 | 7 of 28Sent straight to the model | 23 of 30Sent straight to the model, effort low | 5 of 28Minimal agent with a submission tool, effort xhigh |
| GPT-5.6 Luna | 17 of 28Sent straight to the model | 14 of 30Sent straight to the model, effort max | 21 of 28Minimal agent with a submission tool, effort max |
| GPT-5.6 Sol | 4 of 28Sent straight to the model | 3 of 30Sent straight to the model, effort max | 3 of 28Minimal agent with a submission tool, effort max |
| GPT-5.6 Terra | 6 of 28Sent straight to the model | 5 of 30Sent straight to the model, effort max | 11 of 28Minimal agent with a submission tool, effort max |
| GPT-6 Astra | 1 of 28Sent straight to the model | 1 of 30Sent straight to the model, effort max | 1 of 28Minimal agent with a submission tool, effort max |
The tests here put 224 pairs of models in a different order.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
Show all 224 pairs placed in a different order
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high.
- In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
- In Chess Puzzles, Claude Fable 5, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Fable 5, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Claude Fable 5, sent straight to the model; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5, sent straight to the model, effort high; in Mystery Game Puzzles, Claude Fable 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Claude Fable 5.1, minimal agent with a submission tool, effort max.
- In Chess Puzzles, Claude Fable 5.1, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Claude Fable 5.1, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.6 Sol, sent straight to the model; in Chess Puzzles, GPT-5.6 Sol, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In LiveBench, Claude Fable 5.1, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Fable 5.1, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Fable 5.1, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In Chess Puzzles, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.1, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.5, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Claude Opus 4.5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.1, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.6, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, Claude Opus 4.6, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Opus 4.6, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.1, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.7, sent straight to the model; in Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Claude Opus 4.7, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.7, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.7, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.7, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Opus 4.7, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
- In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in Chess Puzzles, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Claude Opus 4.8, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Opus 4.8, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 4.8, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Claude Opus 4.8, sent straight to the model; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 4.8, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 4.8, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
- In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Claude Opus 5, sent straight to the model; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Sol, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Sol, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Sol, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Opus 5, sent straight to the model, placed above GPT-5.6 Terra, sent straight to the model; in Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Claude Opus 5, sent straight to the model, effort max; in Mystery Game Puzzles, Claude Opus 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, Claude Sonnet 4.6, sent straight to the model, thinking budget 32K, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.5, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash-Lite, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Mystery Game Puzzles, GPT-5.2, minimal agent with a submission tool, effort high, placed above Claude Sonnet 4.6, minimal agent with a submission tool, effort low.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In LiveBench, Claude Sonnet 4.6, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Claude Sonnet 4.6, sent straight to the model, thinking budget 32K; in Mystery Game Puzzles, Claude Sonnet 4.6, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Claude Sonnet 5, sent straight to the model; in Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Claude Sonnet 5, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Claude Sonnet 5, minimal agent with a submission tool, effort max.
- In LiveBench, Claude Sonnet 5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above Claude Sonnet 5, sent straight to the model, effort xhigh; in Mystery Game Puzzles, Claude Sonnet 5, minimal agent with a submission tool, effort max, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3 Flash (preview), sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3 Flash (preview), minimal agent with a submission tool, effort low, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In Chess Puzzles, Gemini 3 Flash (preview), sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3 Flash (preview), minimal agent with a submission tool, effort low.
- In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
- In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.4, sent straight to the model, effort xhigh.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Mystery Game Puzzles, Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.1 Pro (preview), sent straight to the model; in Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Terra, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.1 Pro (preview), sent straight to the model, effort default, placed above GPT-5.6 Terra, sent straight to the model, effort max; in Mystery Game Puzzles, GPT-5.6 Terra, minimal agent with a submission tool, effort max, placed above Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high.
- In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
- In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.2, sent straight to the model, effort xhigh.
- In LiveBench, GPT-5.2, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.5 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.5 Flash, sent straight to the model, effort high, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.5 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In LiveBench, GPT-5.4 Mini, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In Chess Puzzles, GPT-5.4 Mini, sent straight to the model, effort xhigh, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Mini, minimal agent with a submission tool, effort none.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above Gemini 3.5 Flash-Lite, sent straight to the model; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above Gemini 3.5 Flash-Lite, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, Gemini 3.6 Flash, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.6 Flash, sent straight to the model, effort low.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.6 Flash, sent straight to the model, effort low; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.6 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.6 Flash, sent straight to the model, effort low, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above Gemini 3.6 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.6 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, Gemini 3.7 Flash, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.4, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.4, sent straight to the model, effort xhigh.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Gemini 3.7 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.7 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.7 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above Gemini 3.7 Flash, sent straight to the model, effort high; in Mystery Game Puzzles, Gemini 3.7 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Sol, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, effort max.
- In Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Sol, sent straight to the model, effort max; in Mystery Game Puzzles, GPT-5.6 Sol, minimal agent with a submission tool, effort max, placed above Gemini 3.8 Flash, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Chess Puzzles, Gemini 3.8 Flash, sent straight to the model, effort high, placed above GPT-5.6 Terra, sent straight to the model, effort max.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above Gemini 3.8 Flash, sent straight to the model; in Mystery Game Puzzles, Gemini 3.8 Flash, minimal agent with a submission tool, effort high, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.4 Mini, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In Chess Puzzles, GPT-5 mini, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5 mini, minimal agent with a submission tool, effort minimal.
- In Chess Puzzles, GPT-5.1, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.1, minimal agent with a submission tool, effort low.
- In LiveBench, GPT-5.4, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort xhigh.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.2, minimal agent with a submission tool, effort high.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Chess Puzzles, GPT-5.2, sent straight to the model, effort xhigh, placed above GPT-5.6 Luna, sent straight to the model, effort max.
- In LiveBench, GPT-5.6 Luna, sent straight to the model, placed above GPT-5.2, sent straight to the model; in Mystery Game Puzzles, GPT-5.2, minimal agent with a submission tool, effort high, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, GPT-5.4, sent straight to the model, effort xhigh, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.4, minimal agent with a submission tool, effort xhigh.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above GPT-5.4, sent straight to the model; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.4, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.4 Nano, sent straight to the model, placed above GPT-5.4 Mini, sent straight to the model; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.4 Mini, sent straight to the model, effort xhigh; in Mystery Game Puzzles, GPT-5.4 Mini, minimal agent with a submission tool, effort none, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.4 Nano, sent straight to the model; in Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, GPT-5.4 Nano, sent straight to the model, effort high, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.4 Nano, minimal agent with a submission tool, effort none.
- In LiveBench, GPT-5.5, sent straight to the model, placed above GPT-5.6 Luna, sent straight to the model; in Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low.
- In Chess Puzzles, GPT-5.6 Luna, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Luna, minimal agent with a submission tool, effort max.
- In LiveBench, GPT-5.6 Terra, sent straight to the model, placed above GPT-5.5, sent straight to the model; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
- In Chess Puzzles, GPT-5.6 Terra, sent straight to the model, effort max, placed above GPT-5.5, sent straight to the model, effort low; in Mystery Game Puzzles, GPT-5.5, minimal agent with a submission tool, effort xhigh, placed above GPT-5.6 Terra, minimal agent with a submission tool, effort max.
Findings for this category aren't written yet.
LiveBench
best 92.7%LiveBenchNewest result 2026-06-25reliability not measuredcaveat
A test of coding, data analysis, mathematics and reasoning, each graded against a fixed correct answer instead of a judge model. LiveBench refreshes its question sets over time to limit the risk that a model has already seen them. This shows four of LiveBench's seven categories from the 2026-06-25 release; agentic coding uses a different, tool-using harness and is not included here, and language and instruction following have no matching task on this site.
Results
All Sent straight to the model · Run date not published; posted 2026-06-25
Each model's best setting in this test
- GPT-6 Astra92.7%
- Claude Opus 5.592.2%
- Claude Fable 5.191.7%
- GPT-5.6 Sol91.7%
- Claude Opus 591.2%
- GPT-5.6 Terra90.6%
- Claude Fable 589.7%
- GPT-5.589.7%
Show all 29 results (21 not shown above) from LiveBench
- GPT-6 Astra92.7%
- Claude Opus 5.592.2%
- Claude Fable 5.191.7%
- GPT-5.6 Sol91.7%
- Claude Opus 591.2%
- Claude Opus 5.590.7%
- GPT-5.6 Terra90.6%
- Claude Fable 589.7%
- GPT-5.589.7%
- Gemini 3.8 Flash89.3%
- Claude Opus 4.889.2%
- Claude Sonnet 588.7%
- Claude Opus 4.688.7%
- GPT-6 Sol88.7%
- GPT-5.488.1%
- Gemini 3.7 Flash87.8%
- Claude Opus 4.787.2%
- GPT-5.6 Luna85.6%
- Gemini 3.6 Flash85.2%
- Claude Sonnet 4.684.8%
- Gemini 3.1 Pro (preview)84.0%
- GPT-5.283.2%
- Gemini 3.5 Flash82.0%
- GPT-6 Luna81.8%
- GPT-5.4 Nano81.1%
- Claude Opus 4.580.1%
- GPT-5.2 Codex77.7%
- GPT-5.4 Mini71.3%
- Gemini 3.5 Flash-Lite60.2%
Watch out
LiveBench does not publish when it ran each model; the date shown is when the release was posted. LiveBench regularly refreshes, retires and replaces its questions between releases, and has changed which tasks make up a category and rebuilt its agentic coding scoring twice, so a category score from one release is not comparable to the same category on an older or newer release, even for the same model. A model missing from this release's table has no score here; it is not scored zero.
Reliability not measured by this source.
LiveBench is funded by Abacus.AI.
See LiveBench's resultsNot tested here: 6 models
Claude Haiku 4.5, Claude Sonnet 4.5, Gemini 3 Flash (preview), Gemini 3 Pro (preview), GPT-5 mini, GPT-5.1.
More about this test: LiveBench
What's in the test
Each of the four scores here is the unweighted average of that category's task columns in the 2026-06-25 release, using LiveBench's own task names: Coding averages code_generation and code_completion; Data Analysis averages consecutive_events, tablejoin and tablereformat; Mathematics averages AMPS_Hard, integrals_with_game, math_comp and olympiad; Reasoning averages theory_of_mind, zebra_puzzle, spatial and logic_with_navigation.
Technical details for LiveBench
Category average: Each category score is the plain average of that category's task columns, each graded against a fixed correct answer rather than a judge model.
- GPT-6 Astra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5.1, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Opus 5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.6 Terra, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Fable 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Gemini 3.8 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Claude Opus 4.8, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Claude Sonnet 5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Claude Opus 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- GPT-6 Sol, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.4, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- Gemini 3.7 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Claude Opus 4.7, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh Effort.
- GPT-5.6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- Gemini 3.6 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Claude Sonnet 4.6, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Medium Effort.
- Gemini 3.1 Pro (preview), sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-5.2, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- Gemini 3.5 Flash, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
- GPT-6 Luna, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: Max Effort.
- GPT-5.4 Nano, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Claude Opus 4.5, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High Effort.
- GPT-5.2 Codex, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. None.
- GPT-5.4 Mini, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: xHigh.
- Gemini 3.5 Flash-Lite, sent straight to the model: Direct model call. Tools: None. Input: LiveBench's own questions, sent one at a time. Effort: High.
Source version: LiveBench release 2026-06-25
Results posted 2026-06-25
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-27.
Sources for this summary:
Chess Puzzles
best 72.0%Epoch AI100 puzzlesNewest result 2026-09-02reliability not measuredcaveat
Chess positions where the model must find the single best next move, as judged by the Stockfish chess engine. Epoch AI made the puzzles and runs the test itself.
Results
Each model's best setting in this test
- GPT-6 Astra Sent straight to the model, effort max, 2026-08-3072.0%
- Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0261.0%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-1955.0%
- GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0955.0%
- GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0954.0%
- Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2850.0%
- GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1549.0%
- Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-0147.0%
- Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1447.0%
Show all 78 results (69 not shown above) from Chess Puzzles
- GPT-6 Astra Sent straight to the model, effort max, 2026-08-3072.0%
- Gemini 3.8 Flash Sent straight to the model, effort high, 2026-09-0261.0%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort default, 2026-02-1955.0%
- GPT-5.6 Sol Sent straight to the model, effort max, 2026-07-0955.0%
- GPT-5.6 Terra Sent straight to the model, effort max, 2026-07-0954.0%
- Gemini 3.5 Flash Sent straight to the model, effort high, 2026-05-2850.0%
- Gemini 3.1 Pro (preview) Sent straight to the model, effort high, 2026-08-0649.0%
- GPT-5.2 Sent straight to the model, effort xhigh, 2025-12-1549.0%
- Claude Fable 5.1 Sent straight to the model, effort max, 2026-09-0147.0%
- Gemini 3.7 Flash Sent straight to the model, effort high, 2026-08-1447.0%
- Gemini 3.5 Flash Sent straight to the model, effort low, 2026-08-0645.0%
- GPT-5.4 Sent straight to the model, effort xhigh, 2026-03-1144.0%
- Gemini 3.5 Flash Sent straight to the model, effort minimal, 2026-07-1543.0%
- Gemini 3.6 Flash Sent straight to the model, effort low, 2026-08-0743.0%
- Claude Opus 5 Sent straight to the model, effort max, 2026-07-2442.0%
- Claude Fable 5 Sent straight to the model, effort high, 2026-08-0641.0%
- Claude Fable 5 Sent straight to the model, effort max, 2026-06-0941.0%
- Gemini 3 Flash (preview) Sent straight to the model, effort high, 2026-08-0640.0%
- Gemini 3.6 Flash Sent straight to the model, effort high, 2026-08-0240.0%
- GPT-5.2 Sent straight to the model, effort high, 2025-12-1140.0%
- GPT-5.2 Sent straight to the model, effort medium, 2025-12-1140.0%
- GPT-5.6 Luna Sent straight to the model, effort max, 2026-07-0940.0%
- Gemini 3 Flash (preview) Sent straight to the model, effort default, 2025-12-1738.0%
- GPT-5.4 Sent straight to the model, effort high, 2026-07-1538.0%
- GPT-5.4 Sent straight to the model, effort medium, 2026-07-1538.0%
- Claude Sonnet 5 Sent straight to the model, effort xhigh, 2026-06-3035.0%
- Gemini 3.6 Flash Sent straight to the model, effort minimal, 2026-08-0735.0%
- Claude Opus 4.8 Sent straight to the model, effort max, 2026-05-2934.0%
- Claude Opus 5 Sent straight to the model, effort default, 2026-08-0633.0%
- GPT-5.1 Sent straight to the model, effort high, 2025-12-0832.0%
- Gemini 3 Pro (preview) Sent straight to the model, effort default, 2025-12-0831.0%
- Claude Opus 4.7 Sent straight to the model, effort xhigh, 2026-04-2030.0%
- GPT-5 mini Sent straight to the model, effort high, 2026-08-0730.0%
- GPT-5.4 Nano Sent straight to the model, effort high, 2026-04-1430.0%
- Claude Opus 4.8 Sent straight to the model, effort low, 2026-08-0629.0%
- Claude Fable 5 Sent straight to the model, effort low, 2026-08-0628.0%
- GPT-5.6 Sol Sent straight to the model, effort low, 2026-08-0727.0%
- GPT-5.5 Sent straight to the model, effort low, 2026-08-0726.0%
- GPT-5.4 Mini Sent straight to the model, effort xhigh, 2026-08-0724.0%
- GPT-5.2 Sent straight to the model, effort low, 2025-12-1123.0%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort high, 2026-08-0622.0%
- GPT-5.6 Terra Sent straight to the model, effort low, 2026-08-0722.0%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort minimal, 2026-08-0621.0%
- GPT-5.6 Luna Sent straight to the model, effort low, 2026-08-0721.0%
- Claude Opus 4.7 Sent straight to the model, effort low, 2026-07-1420.0%
- Claude Opus 5 Sent straight to the model, effort low, 2026-08-0620.0%
- GPT-5.4 Sent straight to the model, effort low, 2026-07-1520.0%
- Gemini 3.5 Flash-Lite Sent straight to the model, effort low, 2026-08-0618.0%
- GPT-5.4 Mini Sent straight to the model, effort high, 2026-04-1518.0%
- Claude Opus 4.6 Sent straight to the model, thinking budget 32K, 2026-02-0617.0%
- GPT-5.1 Sent straight to the model, effort none, 2026-08-0717.0%
- GPT-5.4 Nano Sent straight to the model, effort low, 2026-08-0717.0%
- Claude Sonnet 5 Sent straight to the model, effort max, 2026-06-3016.0%
- Claude Opus 4.6 Sent straight to the model, effort max, 2026-08-0614.0%
- GPT-5.1 Sent straight to the model, effort low, 2026-08-0714.0%
- Claude Opus 4.6 Sent straight to the model, thinking budget 120K, 2026-02-2013.0%
- Claude Opus 4.8 Sent straight to the model, effort none, 2026-08-0613.0%
- Claude Sonnet 4.6 Sent straight to the model, thinking budget 32K, 2026-02-2013.0%
- Claude Opus 4.5 Sent straight to the model, thinking budget 32K, 2025-12-0812.0%
- Claude Sonnet 4.5 Sent straight to the model, thinking budget 32K, 2025-12-0812.0%
- GPT-5 mini Sent straight to the model, effort low, 2026-08-0712.0%
- Claude Opus 4.6 Sent straight to the model, thinking budget 64K, 2026-02-0610.0%
- GPT-5.5 Sent straight to the model, effort none, 2026-08-0710.0%
- Claude Haiku 4.5 Sent straight to the model, thinking budget 32K, 2026-07-168.0%
- Claude Sonnet 4.6 Sent straight to the model, effort medium, 2026-07-138.0%
- Claude Opus 4.7 Sent straight to the model, effort max, 2026-08-067.0%
- GPT-5 mini Sent straight to the model, effort minimal, 2026-08-077.0%
- GPT-5.6 Sol Sent straight to the model, effort none, 2026-08-077.0%
- Claude Sonnet 4.6 Sent straight to the model, effort high, 2026-07-135.0%
- GPT-5.4 Sent straight to the model, effort none, 2026-07-155.0%
- GPT-5.6 Terra Sent straight to the model, effort none, 2026-08-075.0%
- Claude Opus 4.5 Sent straight to the model, effort default, 2026-08-064.0%
- Claude Sonnet 4.5 Sent straight to the model, effort default, 2026-08-064.0%
- GPT-5.2 Sent straight to the model, effort none, 2026-07-134.0%
- Claude Sonnet 4.6 Sent straight to the model, effort max, 2026-08-063.0%
- GPT-5.4 Mini Sent straight to the model, effort none, 2026-08-073.0%
- GPT-5.4 Nano Sent straight to the model, effort none, 2026-08-073.0%
- GPT-5.6 Luna Sent straight to the model, effort none, 2026-08-072.0%
Watch out
Epoch says chess itself matters little; it uses these puzzles as a rough measure of spatial reasoning and planning. When Epoch tested five human players on 20 puzzles each, the two best got 75% right, more than the best model on the same puzzles. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability not measured by this source.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.
See Epoch AI's resultsNot tested here: 4 models
Claude Opus 5.5, GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.
More about this test: Chess Puzzles
What's in the test
100 new puzzles made by a program, so they do not appear anywhere else. The model gets the board as text in standard chess notation, not as a picture, and answers with one move.
100 puzzles
Technical details for Chess Puzzles
Accuracy: The share of the 100 puzzles where the model chose the best move, in Epoch AI's own runs.
- GPT-6 Astra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.8 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.1 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.6 Sol, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.6 Terra, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.5 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.1 Pro (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Claude Fable 5.1, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3.7 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.5 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Gemini 3.5 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- Gemini 3.6 Flash, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Fable 5, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Fable 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3 Flash (preview), sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3.6 Flash, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.2, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- GPT-5.6 Luna, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Gemini 3 Flash (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.4, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- Claude Sonnet 5, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- Gemini 3.6 Flash, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- Claude Opus 4.8, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Opus 5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.1, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Gemini 3 Pro (preview), sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Claude Opus 4.7, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- GPT-5 mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4 Nano, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Opus 4.8, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Fable 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.6 Sol, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4 Mini, sent straight to the model, effort xhigh: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: xhigh.
- GPT-5.2, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.6 Terra, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- GPT-5.6 Luna, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.7, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 5, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Gemini 3.5 Flash-Lite, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- GPT-5.4 Mini, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- Claude Opus 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- GPT-5.1, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.4 Nano, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Sonnet 5, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- Claude Opus 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.1, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.6, sent straight to the model, thinking budget 120K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 120K tokens.
- Claude Opus 4.8, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Sonnet 4.6, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Claude Opus 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Claude Sonnet 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- GPT-5 mini, sent straight to the model, effort low: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: low.
- Claude Opus 4.6, sent straight to the model, thinking budget 64K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 64K tokens.
- GPT-5.5, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Haiku 4.5, sent straight to the model, thinking budget 32K: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Thinking budget: 32K tokens.
- Claude Sonnet 4.6, sent straight to the model, effort medium: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: medium.
- Claude Opus 4.7, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5 mini, sent straight to the model, effort minimal: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: minimal.
- GPT-5.6 Sol, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Sonnet 4.6, sent straight to the model, effort high: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: high.
- GPT-5.4, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.6 Terra, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Opus 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- Claude Sonnet 4.5, sent straight to the model, effort default: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: default.
- GPT-5.2, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- Claude Sonnet 4.6, sent straight to the model, effort max: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: max.
- GPT-5.4 Mini, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.4 Nano, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
- GPT-5.6 Luna, sent straight to the model, effort none: Epoch AI's evaluation harness (Inspect). Tools: None. Input: Epoch AI's copy of the test questions. Effort: none.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2025-12-08 to 2026-09-02
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
Mystery Game Puzzles
best 84.0%Epoch AI100 puzzlesNewest result 2026-09-02reliability not measuredcaveat
Positions from a puzzle-oriented variant of a well-known game, where the model must pick the single best next move. Epoch AI keeps the game's name secret so that no one can prepare for the test.
Results
Each model's best setting in this test
- GPT-6 Astra Minimal agent with a submission tool, effort max, 2026-08-3084.0%
- Claude Opus 5 Minimal agent with a submission tool, effort max, 2026-07-2559.0%
- Claude Fable 5.1 Minimal agent with a submission tool, effort max, 2026-09-0158.0%
- GPT-5.6 Sol Minimal agent with a submission tool, effort max, 2026-07-2858.0%
- GPT-5.5 Minimal agent with a submission tool, effort xhigh, 2026-07-2456.0%
- Claude Fable 5 Minimal agent with a submission tool, effort max, 2026-07-3152.0%
- Gemini 3.8 Flash Minimal agent with a submission tool, effort high, 2026-09-0247.0%
- Gemini 3.7 Flash Minimal agent with a submission tool, effort high, 2026-08-1437.0%
- GPT-5.4 Minimal agent with a submission tool, effort xhigh, 2026-07-2437.0%
Show all 69 results (60 not shown above) from Mystery Game Puzzles
- GPT-6 Astra Minimal agent with a submission tool, effort max, 2026-08-3084.0%
- Claude Opus 5 Minimal agent with a submission tool, effort max, 2026-07-2559.0%
- Claude Fable 5.1 Minimal agent with a submission tool, effort max, 2026-09-0158.0%
- GPT-5.6 Sol Minimal agent with a submission tool, effort max, 2026-07-2858.0%
- GPT-5.5 Minimal agent with a submission tool, effort xhigh, 2026-07-2456.0%
- Claude Fable 5 Minimal agent with a submission tool, effort max, 2026-07-3152.0%
- GPT-5.5 Minimal agent with a submission tool, effort high, 2026-07-2752.0%
- Gemini 3.8 Flash Minimal agent with a submission tool, effort high, 2026-09-0247.0%
- Claude Opus 5 Minimal agent with a submission tool, effort default, 2026-08-0637.0%
- Gemini 3.7 Flash Minimal agent with a submission tool, effort high, 2026-08-1437.0%
- GPT-5.4 Minimal agent with a submission tool, effort xhigh, 2026-07-2437.0%
- Claude Opus 4.8 Minimal agent with a submission tool, effort max, 2026-07-2536.0%
- Claude Sonnet 5 Minimal agent with a submission tool, effort max, 2026-07-2835.0%
- GPT-5.6 Terra Minimal agent with a submission tool, effort max, 2026-07-2835.0%
- Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort high, 2026-07-2734.0%
- GPT-5.6 Sol Minimal agent with a submission tool, effort none, 2026-08-2733.0%
- Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort medium, 2026-08-0632.0%
- Gemini 3.5 Flash Minimal agent with a submission tool, effort high, 2026-07-2732.0%
- Claude Opus 4.8 Minimal agent with a submission tool, effort xhigh, 2026-07-2631.0%
- Gemini 3.6 Flash Minimal agent with a submission tool, effort high, 2026-08-0530.0%
- Gemini 3.1 Pro (preview) Minimal agent with a submission tool, effort low, 2026-08-0629.0%
- Claude Opus 4.7 Minimal agent with a submission tool, effort max, 2026-07-2528.0%
- Gemini 3.5 Flash Minimal agent with a submission tool, effort low, 2026-08-0528.0%
- GPT-5.4 Minimal agent with a submission tool, effort medium, 2026-08-2828.0%
- GPT-5.5 Minimal agent with a submission tool, effort low, 2026-08-2828.0%
- Gemini 3 Flash (preview) Minimal agent with a submission tool, effort low, 2026-08-0526.0%
- GPT-5.6 Sol Minimal agent with a submission tool, effort low, 2026-08-2726.0%
- Claude Opus 4.6 Minimal agent with a submission tool, effort max, 2026-07-2525.0%
- Gemini 3 Flash (preview) Minimal agent with a submission tool, effort minimal, 2026-08-0525.0%
- Gemini 3.6 Flash Minimal agent with a submission tool, effort minimal, 2026-08-0525.0%
- Gemini 3.6 Flash Minimal agent with a submission tool, effort low, 2026-08-0523.0%
- GPT-5.2 Minimal agent with a submission tool, effort high, 2026-08-0623.0%
- Claude Opus 4.5 Minimal agent with a submission tool, thinking budget 48K, 2026-07-2522.0%
- GPT-5.2 Minimal agent with a submission tool, effort medium, 2026-08-2822.0%
- GPT-5.6 Luna Minimal agent with a submission tool, effort max, 2026-07-2821.0%
- Gemini 3 Flash (preview) Minimal agent with a submission tool, effort high, 2026-08-0520.0%
- GPT-5.6 Luna Minimal agent with a submission tool, effort none, 2026-08-3020.0%
- Gemini 3.5 Flash-Lite Minimal agent with a submission tool, effort low, 2026-08-0519.0%
- GPT-5.1 Minimal agent with a submission tool, effort low, 2026-08-2719.0%
- GPT-5.6 Terra Minimal agent with a submission tool, effort low, 2026-08-2719.0%
- GPT-5.5 Minimal agent with a submission tool, effort none, 2026-08-2718.0%
- Claude Sonnet 4.5 Minimal agent with a submission tool, thinking budget 48K, 2026-07-2517.0%
- GPT-5.4 Minimal agent with a submission tool, effort low, 2026-08-2717.0%
- GPT-5.6 Luna Minimal agent with a submission tool, effort low, 2026-08-2717.0%
- Claude Sonnet 4.6 Minimal agent with a submission tool, effort low, 2026-08-0616.0%
- Claude Sonnet 5 Minimal agent with a submission tool, effort default, 2026-08-0616.0%
- GPT-5.1 Minimal agent with a submission tool, effort medium, 2026-08-2716.0%
- GPT-5.4 Minimal agent with a submission tool, effort none, 2026-08-3016.0%
- Claude Opus 4.6 Minimal agent with a submission tool, effort default, 2026-08-0615.0%
- GPT-5.1 Minimal agent with a submission tool, effort none, 2026-08-2715.0%
- Claude Sonnet 4.6 Minimal agent with a submission tool, effort default, 2026-08-0614.0%
- GPT-5.2 Minimal agent with a submission tool, effort none, 2026-08-2714.0%
- GPT-5.6 Terra Minimal agent with a submission tool, effort none, 2026-08-2714.0%
- Claude Opus 4.7 Minimal agent with a submission tool, effort default, 2026-08-0613.0%
- Gemini 3.5 Flash-Lite Minimal agent with a submission tool, effort minimal, 2026-08-0512.0%
- GPT-5.6 Luna Minimal agent with a submission tool, effort medium, 2026-08-2712.0%
- GPT-5.4 Mini Minimal agent with a submission tool, effort none, 2026-08-2711.0%
- GPT-5 mini Minimal agent with a submission tool, effort minimal, 2026-08-2710.0%
- GPT-5.2 Minimal agent with a submission tool, effort low, 2026-08-2710.0%
- GPT-5.6 Terra Minimal agent with a submission tool, effort medium, 2026-08-2810.0%
- GPT-5.4 Nano Minimal agent with a submission tool, effort none, 2026-08-279.0%
- GPT-5.4 Mini Minimal agent with a submission tool, effort low, 2026-08-278.0%
- Claude Opus 4.6 Minimal agent with a submission tool, effort low, 2026-08-067.0%
- GPT-5.4 Mini Minimal agent with a submission tool, effort medium, 2026-08-277.0%
- GPT-5.4 Nano Minimal agent with a submission tool, effort medium, 2026-08-276.0%
- GPT-5 mini Minimal agent with a submission tool, effort high, 2026-08-275.0%
- GPT-5.4 Nano Minimal agent with a submission tool, effort high, 2026-08-275.0%
- GPT-5 mini Minimal agent with a submission tool, effort medium, 2026-08-274.0%
- GPT-5.4 Nano Minimal agent with a submission tool, effort low, 2026-08-273.0%
Watch out
Epoch does not publish the prompt, example positions or model answers, so readers cannot inspect the questions or answers. An answer with no valid move gets no credit. Epoch AI reports a standard error for its scores on its benchmarking hub. Data: Epoch AI, Capabilities and benchmarking, epoch.ai, CC BY.
Reliability not measured by this source.
Epoch AI is an independent nonprofit supported by donors; its benchmarking is supported by a grant from the UK AI Security Institute.
See Epoch AI's resultsNot tested here: 6 models
Claude Haiku 4.5, Claude Opus 5.5, Gemini 3 Pro (preview), GPT-5.2 Codex, GPT-6 Luna, GPT-6 Sol.
More about this test: Mystery Game Puzzles
What's in the test
100 positions made by a program from random games, kept only where there is a single best move. The model gets the position as text and submits its move with a tool.
100 puzzles
Technical details for Mystery Game Puzzles
Accuracy: The share of the 100 puzzles where the model chose the best move, in Epoch AI's own runs.
- GPT-6 Astra, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Claude Opus 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Claude Fable 5.1, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- GPT-5.6 Sol, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- GPT-5.5, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
- Claude Fable 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- GPT-5.5, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- Gemini 3.8 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- Claude Opus 5, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
- Gemini 3.7 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- GPT-5.4, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
- Claude Opus 4.8, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Claude Sonnet 5, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- GPT-5.6 Terra, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- GPT-5.6 Sol, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- Gemini 3.5 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- Claude Opus 4.8, minimal agent with a submission tool, effort xhigh: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: xhigh.
- Gemini 3.6 Flash, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- Gemini 3.1 Pro (preview), minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Claude Opus 4.7, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Gemini 3.5 Flash, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.4, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.5, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Gemini 3 Flash (preview), minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.6 Sol, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Claude Opus 4.6, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Gemini 3 Flash (preview), minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
- Gemini 3.6 Flash, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
- Gemini 3.6 Flash, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.2, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- Claude Opus 4.5, minimal agent with a submission tool, thinking budget 48K: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Thinking budget: 48K tokens.
- GPT-5.2, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.6 Luna, minimal agent with a submission tool, effort max: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: max.
- Gemini 3 Flash (preview), minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- GPT-5.6 Luna, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.1, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.6 Terra, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.5, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Claude Sonnet 4.5, minimal agent with a submission tool, thinking budget 48K: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Thinking budget: 48K tokens.
- GPT-5.4, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.6 Luna, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Claude Sonnet 4.6, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Claude Sonnet 5, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
- GPT-5.1, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.4, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Claude Opus 4.6, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
- GPT-5.1, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Claude Sonnet 4.6, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
- GPT-5.2, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- GPT-5.6 Terra, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- Claude Opus 4.7, minimal agent with a submission tool, effort default: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: default.
- Gemini 3.5 Flash-Lite, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
- GPT-5.6 Luna, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.4 Mini, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- GPT-5 mini, minimal agent with a submission tool, effort minimal: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: minimal.
- GPT-5.2, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.6 Terra, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.4 Nano, minimal agent with a submission tool, effort none: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: none.
- GPT-5.4 Mini, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- Claude Opus 4.6, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
- GPT-5.4 Mini, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.4 Nano, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5 mini, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- GPT-5.4 Nano, minimal agent with a submission tool, effort high: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: high.
- GPT-5 mini, minimal agent with a submission tool, effort medium: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: medium.
- GPT-5.4 Nano, minimal agent with a submission tool, effort low: Epoch AI's minimal agent setup (Inspect). Tools: A tool for submitting its move. Input: The game position as text. Effort: low.
Source version: Epoch AI benchmark data, downloaded 2026-09-28
Run 2026-07-24 to 2026-09-02
License: Creative Commons Attribution (Epoch AI's own runs). Checked 2026-09-28.
No figures on this page for: Claude Sonnet 5.5, Gemini 2.5 Pro, Gemini 3.1 Flash-Lite (preview), GPT-4.1, GPT-5 nano, GPT-5.1 Codex.
What this doesn't tell you
- Reasoning tests use puzzles and set problems. Doing well on them does not show a model will reason well about your own situation.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- LiveBench: Reliability not measured by this source.
- Chess Puzzles: Reliability not measured by this source.
- Mystery Game Puzzles: Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use