All benchmarks

SWE-bench Pro

Agentic coding

A harder variant of SWE-bench: real bug-fix and feature tasks drawn from actively-maintained repositories, with larger multi-file diffs and no public ground-truth leakage. It measures how reliably a model can resolve genuine GitHub issues end to end.

Model scores

  • Fable 5.181.2%
  • Fable 580.3%
  • Opus 5.589.9%
  • Opus 579.2%
  • Opus 4.869.2%
  • Sonnet 5.581.3%
  • Sonnet 563.2%
  • GPT-6 Astra—
  • GPT-6.1 Sol—
  • GPT-5.6 Sol64.6%
  • Grok 4.7—
  • GPT-5.559.4%
  • Composer 2.5—
  • Opus 4.764.3%
  • Gemini 3.1 Pro54.2%
  • Mythos Preview77.8%

Official source: SWE-bench Pro leaderboard (Scale)

Related reading