Fable 5.1
Fable 5.1
Opus 5.5
Sonnet 5.5
GPT-6.1 Sol
Grok 4.7
vs
GPT-6 Astra
Opus 5.5
Sonnet 5.5
GPT-6 Astra
GPT-6.1 Sol
Grok 4.7
Fable 5.1
$10/$50
per Mtok
Opus 5.5
$4/$20
per Mtok
Sonnet 5.5
$2/$10
per Mtok
GPT-6 Astra
$10/$50
per Mtok
GPT-6.1 Sol
$2/$10
per Mtok
Grok 4.7
$2/$6
per Mtok
Agentic coding
DeepSWE
DeepSWE v1.1
—
74.2%
71.0%
74.1%
75.2%
71.0%
Agentic terminal coding
T-Bench 4.0
Terminal-Bench 4.0
57.9%
64.8%
61.8%
58.2%
58.2%
37.6%
Long context reasoning
AA-LCR v1.1
85.3%
84.7%
82.7%
80.7%
83.0%
76.7%
Agentic financial analysis
Finance Agent v2
58.9%
58.6%
58.1%
53.5%
52.0%
52.3%
Knowledge work
GDPval-AA v2.1
1756
1866
1839
1542
1575
1712
Health
HealthBench Prof.
HealthBench Professional
62.1%
65.6%
69.2%
64.7%
64.2%
56.7%
Agentic coding
SWE-bench Pro
81.2%
89.9%
81.3%
—
—
—
Agentic coding
SWE-bench ML
SWE-bench Multilingual
89.1%
93.9%
90.3%
—
—
—
Agentic coding
FrontierCode 1.1
FrontierCode v1.1 (Main)
50.3%
54.4%
46.2%
53.3%
—
—
Agentic coding
SWE-bench Verified
—
—
—
—
—
—
Multidisciplinary reasoning
HLE (tools)
Humanity's Last Exam (with tools)
65.6%
67.7%
64.5%
57.2%
—
—
Agentic search
BrowseComp
—
—
—
91.5%
—
—
Scaled tool use
MCP-Atlas
—
—
—
—
—
—
Tool use
AutomationBench
AutomationBench 1.0.6 (with fallbacks)
—
42.5%
44.7%
41.4%
36.1%
—
Agentic computer use
OSWorld 2.1
80.7%
partial
42.8%
strict
81.8%
partial
48.7%
strict
80.1%
partial
43.5%
strict
—
—
—
Spatial reasoning
Blueprint-Bench 2
—
—
—
—
—
—
Knowledge work vision
GDPpdf
—
—
—
—
—
—
Legal
Legal Agent
Legal Agent Benchmark
—
8.3%
10.0%
—
—
—
Cybersecurity vulnerability reproduction
CyberGym
—
—
—
—
—
—
Cybersecurity
ExploitBench
ExploitBench (Cap%)
—
91.0%
80.0%
—
—
—
Biology
BioMysteryBench
BioMysteryBench (73/17 split)
—
89.3%
human solvable
50.0%
human difficult
—
—
—
—
Graduate-level reasoning
GPQA Diamond
—
—
—
96.0%
—
—
Visual reasoning
Chartography
44.8%
no tools
88.4%
with tools
64.4%
no tools
89.0%
with tools
61.6%
no tools
90.2%
with tools
—
—
—
Multilingual Q&A
MMMLU
—
—
—
—
—
—
Show 18 more benchmarks
Showing 6 of 24 benchmarks. A dash means the lab has not published a comparable figure.