All benchmarks

SWE-bench Verified

Agentic coding

A 500-problem subset of SWE-bench, each task hand-verified by human engineers as solvable. It is the most widely-cited coding benchmark and a standard proxy for real-world software engineering ability.

Model scores

  • Fable 5.1—
  • Fable 595.0%
  • Opus 5.5—
  • Opus 597.0%
  • Opus 4.888.6%
  • Sonnet 5.5—
  • Sonnet 579.6%
  • GPT-6 Astra—
  • GPT-6.1 Sol—
  • GPT-5.6 Sol96.2%
  • Grok 4.7—
  • GPT-5.582.6%
  • Composer 2.579.6%
  • Opus 4.782.0%
  • Gemini 3.1 Pro78.8%
  • Mythos Preview—

Official source: Vals.ai SWE-bench Verified leaderboard

Related reading