All benchmarks

MCP-Atlas

Scaled tool use

MCP-Atlas evaluates how well a model orchestrates a large catalog of external tools over the Model Context Protocol — selecting the right tool, chaining calls and handling their results across long, multi-tool workflows.

Model scores

  • Fable 5.1—
  • Fable 5—
  • Opus 5.5—
  • Opus 585.8%
  • Opus 4.882.2%
  • Sonnet 5.5—
  • Sonnet 5—
  • GPT-6 Astra—
  • GPT-6.1 Sol—
  • GPT-5.6 Sol—
  • Grok 4.7—
  • GPT-5.575.3%
  • Composer 2.5—
  • Opus 4.779.1%
  • Gemini 3.1 Pro78.2%
  • Mythos Preview—

Official source: MCP-Atlas leaderboard (Scale)

Related reading