All benchmarks

Terminal-Bench 4.0

Agentic terminal coding

Terminal-Bench evaluates models on real-world tasks in a terminal and command-line environment — installing dependencies, debugging, running builds and orchestrating tools — where each step depends on the result of the last. Release 4.0 is the current task set, and the project publishes a single leaderboard on which every vendor runs its own agent harness. Scores are resolution rates with a 95% confidence interval.

Model scores

  • Fable 5.157.9%
  • Fable 544.5%
  • Opus 5.564.8%
  • Opus 553.9%
  • Opus 4.823.6%
  • Sonnet 5.561.8%
  • Sonnet 512.4%
  • GPT-6 Astra58.2%
  • GPT-6.1 Sol58.2%
  • GPT-5.6 Sol37.3%
  • Grok 4.737.6%
  • GPT-5.5—
  • Composer 2.5—
  • Opus 4.7—
  • Gemini 3.1 Pro—
  • Mythos Preview—

Official source: Terminal-Bench 4.0 leaderboard (tbench.ai)

Related reading