Fable 5vsGPT-5.6 Sol
Fable 5
$10/$50
Opus 5
$5/$25
Sonnet 5
$3/$15
GPT-5.6 Sol
$5/$30
Agentic coding
DeepSWE v1.1
69.7%
68.8%
54.0%
72.7%
Agentic coding
SWE-bench Pro
80.3%
79.2%
63.2%
64.6%
Agentic coding
SWE-bench Verified
95.0%
97.0%
79.6%
96.2%
Agentic coding
SWE-bench Multilingual
89.5%
Agentic coding
FrontierCode (Diamond)
29.3%
Long context reasoning
AA-LCR
70.0%
Agentic terminal coding
Terminal-Bench 2.1
88.0%
89.1%
80.4%
88.8%
max
91.9%
ultra
Multidisciplinary reasoning
Humanity's Last Exam
59.0%
no tools
64.5%
with tools
56.3%
no tools
64.7%
with tools
43.2%
no tools
57.4%
with tools
Agentic search
BrowseComp
90.8%
84.7%
90.4%
default
92.2%
ultra
Scaled tool use
MCP-Atlas
85.8%
Tool use
AutomationBench
17.4%
26.0%
13.5%
18.1%
Agentic computer use
OSWorld-Verified
85.0%
81.2%
Spatial reasoning
Blueprint-Bench 2
38.6%
Agentic financial analysis
Finance Agent v2
58.6%
Knowledge work
GDPval-AA
1932
Knowledge work vision
GDPpdf
29.8%
30.7%
Legal
Legal Agent Benchmark
13.3%
11.7%
5.8%
Cybersecurity vulnerability reproduction
CyberGym
83.8%
84.5%
Cybersecurity
ExploitBench (Cap%)
78.0%
70.0%
73.5%
Biology
BioMysteryBench
46.1%
hard
83.9%
human solved
Health
HealthBench Professional
66.0%
59.8%
57.8%
60.5%
Graduate-level reasoning
GPQA Diamond
94.1%
93.7%
94.6%
Visual reasoning
CharXiv Reasoning
77.0%
no tools
88.3%
with tools
Multilingual Q&A
MMMLU