All benchmarks

AA-LCR v1.1

Long context reasoning

Artificial Analysis's Long Context Reasoning evaluation tests whether a model can reason over very large inputs — synthesizing facts scattered across long documents rather than simply retrieving a single passage. Version 1.1 is 100 questions over documents of 10k–100k tokens, graded pass/fail by an LLM judge against an official answer, and run independently by Artificial Analysis on every model.

Model scores

  • Fable 5.185.3%
  • Fable 582.3%
  • Opus 5.584.7%
  • Opus 579.3%
  • Opus 4.877.7%
  • Sonnet 5.582.7%
  • Sonnet 582.0%
  • GPT-6 Astra80.7%
  • GPT-6.1 Sol83.0%
  • GPT-5.6 Sol84.0%
  • Grok 4.776.7%
  • GPT-5.584.3%
  • Composer 2.5—
  • Opus 4.778.7%
  • Gemini 3.1 Pro82.0%
  • Mythos Preview—

Official source: Artificial Analysis — AA-LCR v1.1 leaderboard

Related reading