All benchmarks

Humanity's Last Exam (with tools)

Multidisciplinary reasoning

Humanity's Last Exam (HLE) is a frontier, multi-modal benchmark of expert-authored questions across dozens of academic disciplines, designed to be extremely difficult. This row reports the with-tools configuration, which is the only one every current frontier lab publishes.

Model scores

  • Fable 5.165.6%
  • Fable 563.8%
  • Opus 5.567.7%
  • Opus 563.6%
  • Opus 4.8—
  • Sonnet 5.564.5%
  • Sonnet 554.9%
  • GPT-6 Astra57.2%
  • GPT-6.1 Sol—
  • GPT-5.6 Sol—
  • Grok 4.7—
  • GPT-5.5—
  • Composer 2.5—
  • Opus 4.7—
  • Gemini 3.1 Pro—
  • Mythos Preview—

Official source: Humanity's Last Exam (lastexam.ai)

Related reading