Fable 5.1 vs GPT-6.1 Sol

Fable 5.1 leads, winning 3 of 5 directly-comparable benchmarks against GPT-6.1 Sol.

Head-to-head record: Fable 5.1 3 · 2 GPT-6.1 Sol

Fable 5.1vsGPT-6.1 Sol
Fable 5.1
$10/$50
Fable 5
$10/$50
Opus 5.5
$4/$20
Opus 5
$5/$25
Opus 4.8
$5/$25
Sonnet 5.5
$2/$10
Sonnet 5
$3/$15
GPT-6 Astra
$10/$50
GPT-6.1 Sol
$2/$10
GPT-5.6 Sol
$5/$30
Grok 4.7
$2/$6
GPT-5.5
$5/$30
Composer 2.5
$0.5/$2.5
Opus 4.7
$5/$25
Gemini 3.1 Pro
$2/$12
Mythos Preview
Agentic coding
DeepSWEDeepSWE v1.1
—
69.7%
74.2%
68.8%
59.0%
71.0%
54.0%
74.1%
75.2%
72.7%
71.0%
67.0%
—
—
11.8%
—
Agentic coding
SWE-bench Pro
81.2%
80.3%
89.9%
79.2%
69.2%
81.3%
63.2%
—
—
64.6%
—
59.4%
—
64.3%
54.2%
77.8%
Agentic coding
SWE-bench Verified
—
95.0%
—
97.0%
88.6%
—
79.6%
—
—
96.2%
—
82.6%
79.6%
82.0%
78.8%
—
Agentic coding
SWE-bench MLSWE-bench Multilingual
89.1%
—
93.9%
89.5%
—
90.3%
—
—
—
—
—
77.8%
79.8%
80.5%
—
—
Agentic coding
FrontierCode 1.1FrontierCode v1.1 (Main)
50.3%
—
54.4%
48.0%
—
46.2%
42.4%
53.3%
—
47.5%
—
—
—
—
—
—
Agentic terminal coding
T-Bench 4.0Terminal-Bench 4.0
57.9%
44.5%
64.8%
53.9%
23.6%
61.8%
12.4%
58.2%
58.2%
37.3%
37.6%
—
—
—
—
—
Multidisciplinary reasoning
HLE (tools)Humanity's Last Exam (with tools)
65.6%
63.8%
67.7%
63.6%
—
64.5%
54.9%
57.2%
—
—
—
—
—
—
—
—
Agentic search
BrowseComp
—
—
—
90.8%
84.3%
—
84.7%
—
—
90.4%
default
92.2%
ultra
—
84.4%
—
79.8%
85.9%
86.9%
Scaled tool use
MCP-Atlas
—
—
—
85.8%
82.2%
—
—
—
—
—
—
75.3%
—
79.1%
78.2%
—
Tool use
AutomationBenchAutomationBench 1.0.6 (with fallbacks)
—
—
42.5%
—
—
44.7%
10.7%
41.4%
36.1%
—
—
—
—
—
—
—
Agentic computer use
OSWorld 2.1
80.7%
partial
42.8%
strict
—
81.8%
partial
48.7%
strict
74.0%
partial
37.2%
strict
—
80.1%
partial
43.5%
strict
57.0%
partial
25.6%
strict
—
—
—
—
—
—
—
—
—
Long context reasoning
AA-LCR v1.1
85.3%
82.3%
84.7%
79.3%
77.7%
82.7%
82.0%
80.7%
83.0%
84.0%
76.7%
84.3%
—
78.7%
82.0%
—
Spatial reasoning
Blueprint-Bench 2
—
38.6%
—
—
14.5%
—
—
—
—
—
—
36.2%
—
—
26.5%
—
Agentic financial analysis
Finance Agent v2
58.9%
56.3%
58.6%
58.6%
53.9%
58.1%
53.9%
53.5%
52.0%
53.8%
52.3%
51.8%
—
51.5%
43.0%
—
Knowledge work
GDPval-AA v2.1
1756
1613
1866
1724
1456
1839
1465
1542
1575
1611
1712
1353
—
1356
794
—
Knowledge work vision
GDPpdf
—
29.8%
—
—
22.5%
—
—
—
—
30.7%
—
26.0%
—
—
16.7%
—
Legal
Legal AgentLegal Agent Benchmark
—
13.3%
8.3%
11.7%
10.4%
10.0%
5.8%
—
—
—
—
2.1%
—
—
0.0%
—
Cybersecurity vulnerability reproduction
CyberGym
—
83.8%
—
—
78.8%
—
—
—
—
84.5%
—
81.8%
—
73.1%
—
83.1%
Cybersecurity
ExploitBenchExploitBench (Cap%)
—
78.0%
91.0%
70.0%
40.0%
80.0%
—
—
—
73.5%
—
47.9%
—
—
—
69.0%
Biology
BioMysteryBenchBioMysteryBench (73/17 split)
—
—
89.3%
human solvable
50.0%
human difficult
91.4%
human solvable
51.8%
human difficult
—
—
84.9%
human solvable
39.4%
human difficult
—
—
—
—
—
—
—
—
90.3%
human solvable
44.1%
human difficult
Health
HealthBench Prof.HealthBench Professional
62.1%
66.0%
65.6%
59.8%
56.9%
69.2%
57.8%
64.7%
64.2%
60.5%
56.7%
49.5%
—
—
—
64.7%
Graduate-level reasoning
GPQA Diamond
—
94.1%
—
93.7%
93.6%
—
—
96.0%
—
94.6%
—
93.6%
—
94.2%
94.3%
94.6%
Visual reasoning
Chartography
44.8%
no tools
88.4%
with tools
—
64.4%
no tools
89.0%
with tools
29.8%
no tools
83.4%
with tools
—
61.6%
no tools
90.2%
with tools
—
—
—
—
—
—
—
—
—
—
Multilingual Q&A
MMMLU
—
—
—
—
—
—
—
—
—
—
—
83.2%
—
91.5%
92.6%
—