AA-LCR v1.1
Artificial Analysis's Long Context Reasoning evaluation tests whether a model can reason over very large inputs — synthesizing facts scattered across long documents rather than simply retrieving a single passage. Version 1.1 is 100 questions over documents of 10k–100k tokens, graded pass/fail by an LLM judge against an official answer, and run independently by Artificial Analysis on every model.
Official source: Artificial Analysis — AA-LCR v1.1 leaderboard