Agent reliability
Whether an AI agent gets the same job right every time, not just once. Good means it succeeds on repeated tries.
Updated every Friday. Sources last checked 2026-09-26.
At a glance
Vendor claims have not been collected for any category yet.
ⓘ More about this category
- Reliability figures shown
- 8 sources, 17 models with figures
- Sources last checked 2026-09-29
- Ring 4: the tester says it repeats its runs; its method is under the source.
- BFCL = Berkeley Function-Calling Leaderboard, T2R = tau2-bench retail, T2A = tau2-bench airline, T2T = tau2-bench telecom, T3B = tau3-bench banking, TB = Terminal-Bench, RDY = READY, MCPU = MCP-Universe
Who does well
| Test | Best | Best score | Next | Next score |
|---|---|---|---|---|
| Berkeley Function-Calling Leaderboard | Claude Opus 4.5 Native function calling | 77.5% | Claude Sonnet 4.5 Native function calling | 73.2% |
| tau2-bench retail | GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high | 81.6% | Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high | 79.6% |
| tau2-bench airline | Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort highroughly tied | 84.0% | GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high | 83.0% |
| tau2-bench telecom | Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high | 92.3% | Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high | 91.2% |
| tau3-bench banking | Claude Opus 5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max | 48.7% | GPT-5.6 Sol Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh | 46.9% |
| Model | Berkeley Function-Calling Leaderboard | tau2-bench retail | tau2-bench airline | tau2-bench telecom | tau3-bench banking |
|---|---|---|---|---|---|
| Claude Opus 4.5 | 1 of 6Native function calling | 2 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 1 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 1 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 13 of 15Agent with Sierra's standard search tools, talking to an AI-played customer, effort high |
| Claude Sonnet 4.5 | 2 of 6Native function calling | 5 of 5Agent with the domain's tools, talking to an AI-played customer, extended thinking on | 5 of 5Agent with the domain's tools, talking to an AI-played customer, extended thinking on | 5 of 5Agent with the domain's tools, talking to an AI-played customer, extended thinking on | 12 of 15Agent searching with a terminal only, talking to an AI-played customer, extended thinking on |
| Gemini 3 Flash (preview) | 3 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 3 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 2 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 9 of 15Agent searching with a terminal only, talking to an AI-played customer, effort high | |
| Gemini 3 Pro (preview) | 3 of 6Tools described in the prompt | 4 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 4 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 3 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 14 of 15Agent searching with a terminal only, talking to an AI-played customer, effort high |
| GPT-5.2 | 5 of 6Native function calling | 1 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 2 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 4 of 5Agent with the domain's tools, talking to an AI-played customer, effort high | 8 of 15Agent with Sierra's standard search tools, talking to an AI-played customer, effort high |
The tests here put 32 pairs of models in a different order.
- In Berkeley Function-Calling Leaderboard, Claude Opus 4.5, native function calling, placed above Claude Sonnet 4.5, native function calling; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench retail, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
Show all 32 pairs placed in a different order
- In Berkeley Function-Calling Leaderboard, Claude Opus 4.5, native function calling, placed above Claude Sonnet 4.5, native function calling; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench retail, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench retail, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, Gemini 3 Flash (preview), agent searching with a terminal only, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, Gemini 3 Flash (preview), agent searching with a terminal only, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, Gemini 3 Flash (preview), agent searching with a terminal only, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Claude Opus 4.5, native function calling, placed above GPT-5.2, native function calling; in tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Claude Opus 4.5, native function calling, placed above GPT-5.2, native function calling; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench airline, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench telecom, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above Gemini 3 Pro (preview), tools described in the prompt; in tau2-bench retail, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above Gemini 3 Pro (preview), tools described in the prompt; in tau2-bench airline, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above Gemini 3 Pro (preview), tools described in the prompt; in tau2-bench telecom, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In tau2-bench retail, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
- In tau2-bench airline, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on; in tau3-bench banking, Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on, placed above Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above GPT-5.2, native function calling; in tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above GPT-5.2, native function calling; in tau2-bench airline, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above GPT-5.2, native function calling; in tau2-bench telecom, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on.
- In Berkeley Function-Calling Leaderboard, Claude Sonnet 4.5, native function calling, placed above GPT-5.2, native function calling; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on.
- In tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench telecom, Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench telecom, Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Gemini 3 Flash (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Gemini 3 Pro (preview), tools described in the prompt, placed above GPT-5.2, native function calling; in tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Gemini 3 Pro (preview), tools described in the prompt, placed above GPT-5.2, native function calling; in tau2-bench airline, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high.
- In Berkeley Function-Calling Leaderboard, Gemini 3 Pro (preview), tools described in the prompt, placed above GPT-5.2, native function calling; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
- In tau2-bench retail, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench telecom, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench airline, GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high; in tau2-bench telecom, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high.
- In tau2-bench telecom, Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high, placed above GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high; in tau3-bench banking, GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high, placed above Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high.
Quality and cost
In Berkeley Function-Calling Leaderboard only. Cost as published by UC Berkeley Gorilla project. Your costs will differ.
Most accurate, and the cheapest within 1 point of the best: Claude Opus 4.5 Native function calling 77.5% · $86.55 for BFCL's full run
Best at each price
- 1Claude Haiku 4.5 Native function calling 68.7% · $14.23 for BFCL's full run
- 2Claude Sonnet 4.5 Native function calling 73.2% · $43.73 for BFCL's full run
- 3Claude Opus 4.5 Native function calling 77.5% · $86.55 for BFCL's full runbest quality
Findings for this category aren't written yet.
Berkeley Function-Calling Leaderboard
best 77.5%UC Berkeley Gorilla projectNewest result 2025-12-16reliability not measuredcaveat
BFCL, the Berkeley Function-Calling Leaderboard, tests whether a model calls the right tool with the right arguments. It covers single tool calls and calls that must be made in sequence or in parallel. Other tests cover conversations where the model has to call tools over several turns, agentic web search and agentic memory. It also checks whether the model correctly stays quiet when no tool fits the request.
Results
Run date not published; posted 2025-12-16
Each model's best setting in this test
- Claude Opus 4.5 Native function calling77.5%
- Claude Sonnet 4.5 Native function calling73.2%
- Gemini 3 Pro (preview) Tools described in the prompt72.5%
- Claude Haiku 4.5 Native function calling68.7%
- GPT-5.2 Native function calling55.9%
- GPT-5 mini Native function calling55.5%
Show all 12 results (6 not shown above) from Berkeley Function-Calling Leaderboard
- Claude Opus 4.5 Native function calling77.5%
- Claude Sonnet 4.5 Native function calling73.2%
- Gemini 3 Pro (preview) Tools described in the prompt72.5%
- Claude Haiku 4.5 Native function calling68.7%
- Gemini 3 Pro (preview) Native function calling68.1%
- GPT-5.2 Native function calling55.9%
- GPT-5 mini Native function calling55.5%
- GPT-5.2 Tools described in the prompt45.3%
- Claude Opus 4.5 Tools described in the prompt33.5%
- GPT-5 mini Tools described in the prompt27.8%
- Claude Haiku 4.5 Tools described in the prompt25.3%
- Claude Sonnet 4.5 Tools described in the prompt24.9%
Watch out
BFCL does not publish when it ran each model; the date shown is when BFCL updated its leaderboard with these results. A model can appear here twice, once run with its own tool-calling interface (FC) and once with the tools only described in the prompt (Prompt); these are separate setups and are never combined. The published cost is BFCL's own estimate for its full run of that model, not a per-question figure. Data: Berkeley Function-Calling Leaderboard (BFCL), Apache 2.0.
Reliability not measured by this source.
BFCL is run by UC Berkeley's Gorilla project.
See UC Berkeley Gorilla project's resultsNot tested here: 11 models
Claude Fable 5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Gemini 2.5 Pro, Gemini 3 Flash (preview), Gemini 3.1 Pro (preview), GPT-5.4, GPT-5.5, GPT-5.6 Sol.
More about this test: Berkeley Function-Calling Leaderboard
What's in the test
BFCL can run a model two ways. FC gives the model its own tool-calling interface. Prompt only describes the tools in the prompt text and has the model write out the call the normal way it writes text. A model can appear on this page twice, once under each setup, and the two are never combined.
- Non-live function calls
- Live (user-contributed) function calls
- Multi-turn conversations
- Agentic web search
- Agentic memory
- Relevance and irrelevance detection
Technical details for Berkeley Function-Calling Leaderboard
Overall accuracy: Overall Acc is BFCL's own combination of its category scores within one test: non-live and live function calls, multi-turn conversations, web search, memory, and telling correctly when no tool applies. It is an unweighted average of those categories, not a plain share of individual test cases answered correctly.
- Claude Opus 4.5, native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- Claude Sonnet 4.5, native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- Gemini 3 Pro (preview), tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
- Claude Haiku 4.5, native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- Gemini 3 Pro (preview), native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- GPT-5.2, native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- GPT-5 mini, native function calling: BFCL, native function calling. Tools: The model's own tool-calling interface. Input: BFCL's own test cases. None.
- GPT-5.2, tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
- Claude Opus 4.5, tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
- GPT-5 mini, tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
- Claude Haiku 4.5, tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
- Claude Sonnet 4.5, tools described in the prompt: BFCL, tools described in the prompt. Tools: Tool definitions in the prompt text. Input: BFCL's own test cases. None.
Source version: BFCL results, 2025-12-16
Results posted 2025-12-16
Run dates not published by the tester; the dates are when results were posted.
License: Apache License 2.0. Checked 2026-09-28.
tau2-bench retail
best 81.6%Sierra115 tasksNewest result 2026-03-02reliability figures showncaveat
tau2-bench, built by Sierra, puts an AI agent in a customer-service job: it talks with a customer, played by another AI model, and must follow the business's policy while using the business's tools to finish the task. This source is the retail domain: orders, returns and exchanges.
Results
Each model's best setting in this test
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2681.6%
Completes it 4 tries out of 4: 51.8%
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2679.6%
Completes it 4 tries out of 4: 51.8%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0276.8%
Completes it 4 tries out of 4: 51.8%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0275.9%
Completes it 4 tries out of 4: 47.4%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2672.4%
Completes it 4 tries out of 4: 39.5%
Show all 6 results (1 not shown above) from tau2-bench retail
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2681.6%
Completes it 4 tries out of 4: 51.8%
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2679.6%
Completes it 4 tries out of 4: 51.8%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0276.8%
Completes it 4 tries out of 4: 51.8%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0275.9%
Completes it 4 tries out of 4: 47.4%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort none, Run date not published; posted 2026-02-2675.0%
Completes it 4 tries out of 4: 45.6%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2672.4%
Completes it 4 tries out of 4: 39.5%
Watch out
Sierra, which runs this test, sells AI customer-service agents. Sierra does not publish when it ran each model; the date shown is when the results were posted. The newest models on Sierra's site were run only on a separate banking domain, shown as its own test on this page. Each domain (retail, airline, telecom) is a different set of tasks, so compare models within one domain, never across them. Data: Sierra, tau-bench leaderboard, shown with Sierra's permission.
Reliability: Sierra runs each task more than once. The reliability figure is pass^4: the chance the model completes the same task on 4 tries out of 4. A big gap between the two numbers means the model is inconsistent.
Sierra built tau-bench and runs every result shown here. Sierra sells AI customer-service agents for retail, airline and telecom, the same kinds of tasks this test covers. It does not make any of the models shown.
See Sierra's resultsNot tested here: 12 models
Claude Fable 5, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Gemini 2.5 Pro, Gemini 3.1 Pro (preview), GPT-5 mini, GPT-5.4, GPT-5.5, GPT-5.6 Sol.
More about this test: tau2-bench retail
What's in the test
Each task is a customer conversation that ends with the store's records either matching the right outcome or not. Sierra runs each task more than once, so it can report both how often the agent succeeds and how often it succeeds every time.
115 tasks
- Customer conversations with tool use
- Policy following
Technical details for tau2-bench retail
Success rate (pass^1): The share of the 115 retail customer-service tasks the agent completed, averaged over Sierra's repeated runs of each task.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort none: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: none.
- Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Extended thinking on.
Source version: Sierra tau-bench submissions, gpt-5-2_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, claude-opus-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-flash_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-pro_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2-none_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, claude-sonnet-4-5_sierra_2026-02-26, downloaded 2026-09-29
Results posted 2026-02-26 to 2026-03-02
Run dates not published by the tester; the dates are when results were posted.
License: Display permitted by Sierra in writing, 2026-09-28. Checked 2026-09-29.
tau2-bench airline
best 84.0%Sierra50 tasksNewest result 2026-03-02reliability figures showncaveat
tau2-bench, built by Sierra, puts an AI agent in a customer-service job: it talks with a customer, played by another AI model, and must follow the business's policy while using the business's tools to finish the task. This source is the airline domain: bookings, changes and cancellations.
Results
Each model's best setting in this test
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2684.0%
Completes it 4 tries out of 4: 70.0%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2683.0%
Completes it 4 tries out of 4: 72.0%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0282.5%
Completes it 4 tries out of 4: 68.0%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0280.5%
Completes it 4 tries out of 4: 66.0%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2672.0%
Completes it 4 tries out of 4: 48.0%
Show all 6 results (1 not shown above) from tau2-bench airline
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2684.0%
Completes it 4 tries out of 4: 70.0%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2683.0%
Completes it 4 tries out of 4: 72.0%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0282.5%
Completes it 4 tries out of 4: 68.0%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0280.5%
Completes it 4 tries out of 4: 66.0%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2672.0%
Completes it 4 tries out of 4: 48.0%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort none, Run date not published; posted 2026-02-2652.5%
Completes it 4 tries out of 4: 22.0%
Watch out
Sierra, which runs this test, sells AI customer-service agents. Sierra does not publish when it ran each model; the date shown is when the results were posted. The newest models on Sierra's site were run only on a separate banking domain, shown as its own test on this page. Each domain (retail, airline, telecom) is a different set of tasks, so compare models within one domain, never across them. Data: Sierra, tau-bench leaderboard, shown with Sierra's permission.
Reliability: Sierra runs each task more than once. The reliability figure is pass^4: the chance the model completes the same task on 4 tries out of 4. A big gap between the two numbers means the model is inconsistent.
Sierra built tau-bench and runs every result shown here. Sierra sells AI customer-service agents for retail, airline and telecom, the same kinds of tasks this test covers. It does not make any of the models shown.
See Sierra's resultsNot tested here: 12 models
Claude Fable 5, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Gemini 2.5 Pro, Gemini 3.1 Pro (preview), GPT-5 mini, GPT-5.4, GPT-5.5, GPT-5.6 Sol.
More about this test: tau2-bench airline
What's in the test
Each task is a customer conversation that ends with the business's records either matching the right outcome or not. Sierra runs each task more than once, so it can report both how often the agent succeeds and how often it succeeds every time.
50 tasks
- Customer conversations with tool use
- Policy following
Technical details for tau2-bench airline
Success rate (pass^1): The share of the 50 airline customer-service tasks the agent completed, averaged over Sierra's repeated runs of each task.
- Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Extended thinking on.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort none: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: none.
Source version: Sierra tau-bench submissions, claude-opus-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-flash_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-pro_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, claude-sonnet-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2-none_sierra_2026-02-26, downloaded 2026-09-29
Results posted 2026-02-26 to 2026-03-02
Run dates not published by the tester; the dates are when results were posted.
License: Display permitted by Sierra in writing, 2026-09-28. Checked 2026-09-29.
tau2-bench telecom
best 92.3%Sierra114 tasksNewest result 2026-03-02reliability figures showncaveat
tau2-bench, built by Sierra, puts an AI agent in a customer-service job: it talks with a customer, played by another AI model, and must follow the business's policy while using the business's tools to finish the task. This source is the telecom domain: troubleshooting a customer's phone service, where the customer also has to take actions on their own device.
Results
Each model's best setting in this test
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2692.3%
Completes it 4 tries out of 4: 78.1%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0291.2%
Completes it 4 tries out of 4: 70.2%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0291.0%
Completes it 4 tries out of 4: 74.6%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2689.7%
Completes it 4 tries out of 4: 71.9%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2684.9%
Completes it 4 tries out of 4: 64.0%
Show all 6 results (1 not shown above) from tau2-bench telecom
- Claude Opus 4.5 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2692.3%
Completes it 4 tries out of 4: 78.1%
- Gemini 3 Flash (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0291.2%
Completes it 4 tries out of 4: 70.2%
- Gemini 3 Pro (preview) Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-03-0291.0%
Completes it 4 tries out of 4: 74.6%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort high, Run date not published; posted 2026-02-2689.7%
Completes it 4 tries out of 4: 71.9%
- Claude Sonnet 4.5 Agent with the domain's tools, talking to an AI-played customer, extended thinking on, Run date not published; posted 2026-02-2684.9%
Completes it 4 tries out of 4: 64.0%
- GPT-5.2 Agent with the domain's tools, talking to an AI-played customer, effort none, Run date not published; posted 2026-02-2657.2%
Completes it 4 tries out of 4: 30.7%
Watch out
Sierra, which runs this test, sells AI customer-service agents. Sierra does not publish when it ran each model; the date shown is when the results were posted. The newest models on Sierra's site were run only on a separate banking domain, shown as its own test on this page. Each domain (retail, airline, telecom) is a different set of tasks, so compare models within one domain, never across them. Data: Sierra, tau-bench leaderboard, shown with Sierra's permission.
Reliability: Sierra runs each task more than once. The reliability figure is pass^4: the chance the model completes the same task on 4 tries out of 4. A big gap between the two numbers means the model is inconsistent.
Sierra built tau-bench and runs every result shown here. Sierra sells AI customer-service agents for retail, airline and telecom, the same kinds of tasks this test covers. It does not make any of the models shown.
See Sierra's resultsNot tested here: 12 models
Claude Fable 5, Claude Haiku 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Gemini 2.5 Pro, Gemini 3.1 Pro (preview), GPT-5 mini, GPT-5.4, GPT-5.5, GPT-5.6 Sol.
More about this test: tau2-bench telecom
What's in the test
Each task is a customer conversation that ends with the business's records either matching the right outcome or not. Sierra runs each task more than once, so it can report both how often the agent succeeds and how often it succeeds every time.
114 tasks
- Customer conversations with tool use
- Policy following
Technical details for tau2-bench telecom
Success rate (pass^1): The share of the 114 telecom customer-service tasks the agent completed, averaged over Sierra's repeated runs of each task.
- Claude Opus 4.5, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Flash (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Pro (preview), agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: high.
- Claude Sonnet 4.5, agent with the domain's tools, talking to an AI-played customer, extended thinking on: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Extended thinking on.
- GPT-5.2, agent with the domain's tools, talking to an AI-played customer, effort none: Standard tau2-bench setup, run by Sierra. Tools: The domain's tools and policy guidelines. Input: A conversation with a customer played by an AI model. Effort: none.
Source version: Sierra tau-bench submissions, claude-opus-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-flash_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-pro_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, claude-sonnet-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2-none_sierra_2026-02-26, downloaded 2026-09-29
Results posted 2026-02-26 to 2026-03-02
Run dates not published by the tester; the dates are when results were posted.
License: Display permitted by Sierra in writing, 2026-09-28. Checked 2026-09-29.
tau3-bench banking
best 48.7%Sierra97 tasksNewest result 2026-08-03reliability figures showncaveat
tau3-bench banking, built by Sierra and published as tau-knowledge, puts an AI agent in a bank's customer-service job. It talks with a customer, played by another AI model, and must find what it needs in a knowledge base of about 700 bank documents, follow the bank's policy and use the bank's tools to finish the task.
Results
Each model's best setting in this test
- Claude Opus 5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-08-0348.7%
Completes it 4 tries out of 4: 32.0%
- GPT-5.6 Sol Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-07-2246.9%
Completes it 4 tries out of 4: 27.8%
- GPT-5.5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-07-2244.6%
Completes it 4 tries out of 4: 29.9%
- Claude Opus 4.7 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2340.2%
Completes it 4 tries out of 4: 24.7%
- Claude Fable 5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2339.7%
Completes it 4 tries out of 4: 28.9%
- Claude Opus 4.8 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2339.7%
Completes it 4 tries out of 4: 22.7%
- GPT-5.4 Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-05-0639.4%
Completes it 4 tries out of 4: 21.6%
- GPT-5.2 Agent with Sierra's standard search tools, talking to an AI-played customer, effort high, 2026-05-0532.2%
Completes it 4 tries out of 4: 18.6%
Show all 16 results (8 not shown above) from tau3-bench banking
- Claude Opus 5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-08-0348.7%
Completes it 4 tries out of 4: 32.0%
- GPT-5.6 Sol Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-07-2246.9%
Completes it 4 tries out of 4: 27.8%
- GPT-5.5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-07-2244.6%
Completes it 4 tries out of 4: 29.9%
- Claude Opus 4.7 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2340.2%
Completes it 4 tries out of 4: 24.7%
- Claude Fable 5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2339.7%
Completes it 4 tries out of 4: 28.9%
- Claude Opus 4.8 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-07-2339.7%
Completes it 4 tries out of 4: 22.7%
- GPT-5.4 Agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh, 2026-05-0639.4%
Completes it 4 tries out of 4: 21.6%
- GPT-5.2 Agent with Sierra's standard search tools, talking to an AI-played customer, effort high, 2026-05-0532.2%
Completes it 4 tries out of 4: 18.6%
- Claude Opus 4.6 Agent with Sierra's standard search tools, talking to an AI-played customer, effort max, 2026-05-0627.3%
Completes it 4 tries out of 4: 11.3%
- Gemini 3 Flash (preview) Agent searching with a terminal only, talking to an AI-played customer, effort high, 2026-03-0227.3%
Completes it 4 tries out of 4: 7.2%
- Gemini 3.1 Pro (preview) Agent with Sierra's standard search tools, talking to an AI-played customer, effort high, 2026-05-0526.0%
Completes it 4 tries out of 4: 9.3%
- Claude Sonnet 4.5 Agent searching with a terminal only, talking to an AI-played customer, extended thinking on, 2026-02-2425.3%
Completes it 4 tries out of 4: 10.3%
- Claude Opus 4.5 Agent with Sierra's standard search tools, talking to an AI-played customer, effort high, 2026-05-0524.7%
Completes it 4 tries out of 4: 11.3%
- Gemini 3 Pro (preview) Agent searching with a terminal only, talking to an AI-played customer, effort high, 2026-03-0218.0%
Completes it 4 tries out of 4: 4.1%
- Gemini 2.5 Pro Agent with Sierra's standard search tools, talking to an AI-played customer, effort high, 2026-05-0513.7%
Completes it 4 tries out of 4: 1.0%
- GPT-5.2 Agent searching with Qwen embeddings only, talking to an AI-played customer, effort none, 2026-05-2612.6%
Completes it 4 tries out of 4: 4.1%
Watch out
Sierra, which runs this test, sells AI customer-service agents. The date shown is the day Sierra ran each model. Most rows let the agent search the documents with Sierra's standard set of search tools; four rows used a different search method, named in each row's setup, so compare those with care. Sierra re-scored its older banking runs on 2026-07-15 after fixing its grading and says earlier scores are not comparable; only results scored with the fixed grading are shown. For Gemini 3 Flash and Gemini 3 Pro, Sierra reports that 52 and 42 runs were stopped at the step limit and counted as failures. Banking is a different set of tasks from retail, airline and telecom, so compare models within one test, never across them. Data: Sierra, tau-bench leaderboard, shown with Sierra's permission.
Reliability: Sierra runs each task more than once. The reliability figure is pass^4: the chance the model completes the same task on 4 tries out of 4. A big gap between the two numbers means the model is inconsistent.
Sierra built tau-bench and runs every result shown here. Sierra sells AI customer-service agents, the same kind of work this test covers. It does not make any of the models shown.
See Sierra's resultsNot tested here: 2 models
Claude Haiku 4.5, GPT-5 mini.
More about this test: tau3-bench banking
What's in the test
Each task is a customer conversation that ends with the bank's records either matching the right outcome or not. The agent has to search the bank's documents to know what to do. Sierra runs each task four times, so it can report both how often the agent succeeds and how often it succeeds every time.
97 tasks
- Customer conversations with tool use
- Searching a knowledge base
- Policy following
Technical details for tau3-bench banking
Success rate (pass^1): The share of the 97 banking customer-service tasks the agent completed, averaged over Sierra's repeated runs of each task.
- Claude Opus 5, agent with Sierra's standard search tools, talking to an AI-played customer, effort max: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: max.
- GPT-5.6 Sol, agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: xhigh.
- GPT-5.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: xhigh.
- Claude Opus 4.7, agent with Sierra's standard search tools, talking to an AI-played customer, effort max: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: max.
- Claude Fable 5, agent with Sierra's standard search tools, talking to an AI-played customer, effort max: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: max.
- Claude Opus 4.8, agent with Sierra's standard search tools, talking to an AI-played customer, effort max: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: max.
- GPT-5.4, agent with Sierra's standard search tools, talking to an AI-played customer, effort xhigh: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: xhigh.
- GPT-5.2, agent with Sierra's standard search tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- Claude Opus 4.6, agent with Sierra's standard search tools, talking to an AI-played customer, effort max: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: max.
- Gemini 3 Flash (preview), agent searching with a terminal only, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: A terminal to search the bank's documents (not Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3.1 Pro (preview), agent with Sierra's standard search tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- Claude Sonnet 4.5, agent searching with a terminal only, talking to an AI-played customer, extended thinking on: Standard tau2-bench setup, run by Sierra. Tools: A terminal to search the bank's documents (not Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Extended thinking on.
- Claude Opus 4.5, agent with Sierra's standard search tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 3 Pro (preview), agent searching with a terminal only, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: A terminal to search the bank's documents (not Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- Gemini 2.5 Pro, agent with Sierra's standard search tools, talking to an AI-played customer, effort high: Standard tau2-bench setup, run by Sierra. Tools: Keyword search, meaning search and a sandboxed shell over the bank's documents (Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: high.
- GPT-5.2, agent searching with Qwen embeddings only, talking to an AI-played customer, effort none: Standard tau2-bench setup, run by Sierra. Tools: Meaning search with Qwen embeddings over the bank's documents (not Sierra's standard setting), plus the bank's tools and policy. Input: A conversation with a customer played by an AI model. Effort: none.
Source version: Sierra tau-bench submissions, claude-opus-5_sierra_2026-08-04, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-6-sol_sierra_2026-08-04, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-5_sierra_2026-05-05, downloaded 2026-09-29; Sierra tau-bench submissions, claude-opus-4-7_sierra_2026-05-05, downloaded 2026-09-29; Sierra tau-bench submissions, claude-fable-5_sierra_2026-08-04, downloaded 2026-09-29; Sierra tau-bench submissions, claude-opus-4-8_sierra_2026-08-04, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-4_sierra_2026-03-25, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, claude-opus-4-6_sierra_2026-05-05, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-flash_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-1-pro-preview_sierra_2026-05-05, downloaded 2026-09-29; Sierra tau-bench submissions, claude-sonnet-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, claude-opus-4-5_sierra_2026-02-26, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-3-pro_sierra_2026-03-02, downloaded 2026-09-29; Sierra tau-bench submissions, gemini-2-5-pro_sierra_2026-05-05, downloaded 2026-09-29; Sierra tau-bench submissions, gpt-5-2-none_sierra_2026-02-26, downloaded 2026-09-29
Run 2026-02-24 to 2026-08-03
License: Display permitted by Sierra in writing, 2026-09-28. Checked 2026-09-29.
No figures on this page for: Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5, Gemini 3.1 Flash-Lite (preview), Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, GPT-4.1, GPT-5 nano, GPT-5.1, GPT-5.1 Codex, GPT-5.2 Codex, GPT-5.4 Mini, GPT-5.4 Nano, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol.
Also tested by
Their figures are not shown here because permission to reuse them is not yet confirmed.
Terminal-Bench
Terminal-Bench (Stanford and Laude Institute), checked 2026-09-24The share of Terminal-Bench command-line tasks an agent resolves, run as repeated trials per task and shown with a 95% confidence interval.
Last updated 2026-09-21
See their resultsMore about this test: Terminal-Bench
Reliability: Terminal-Bench runs repeated trials of each task and reports the uncertainty of each resolution rate as a 95% confidence interval.
Technical details for Terminal-Bench
Resolution rate: The share of Terminal-Bench command-line tasks an agent resolves, run as repeated trials per task and shown with a 95% confidence interval.
License: Results license not stated. Checked 2026-09-24.
READY
Scale AI, checked 2026-09-24The share of held-out cases an agent completes correctly once its human-oversight policy has been tuned to hit a set target. READY calls this reliability; it is a success rate across different cases, not a repeat of the same case.
Scale AI uses a held-out qualification step to show how much human review an agent needs to meet a chosen reliability target.
Last updated: not published by the tester
See their resultsMore about this test: READY
Reliability not measured by this source.
Technical details for READY
Reliability: The share of held-out cases an agent completes correctly once its human-oversight policy has been tuned to hit a set target. READY calls this reliability; it is a success rate across different cases, not a repeat of the same case.
License: Scale website terms (internal use only). Checked 2026-09-24.
MCP-Universe
Salesforce AI Research, checked 2026-09-28The share of MCP-Universe tasks an agent completes using real tool servers (MCP), as graded by Salesforce.
Last updated 2026-01-26
See their resultsMore about this test: MCP-Universe
Reliability not measured by this source.
Salesforce AI Research built MCP-Universe and runs and grades every result. Salesforce sells AI agents (Agentforce). The leaderboard was last updated in January 2026.
Technical details for MCP-Universe
Success rate: The share of MCP-Universe tasks an agent completes using real tool servers (MCP), as graded by Salesforce.
License: Linked only; see the leaderboard. Checked 2026-09-28.
Sources for this summary:
What this doesn't tell you
- Few testers measure consistency across repeated runs, so the evidence here is thin.
- None of these tests used your customers or your systems.
- Scores from different sources are not comparable, even when they share a unit.
- None of these tests were run on your own documents, code or data.
- Berkeley Function-Calling Leaderboard: Reliability not measured by this source.
Data version 2026-09-30+832fb354fcbe · Terms of use