All benchmarks

OSWorld 2.1

Agentic computer use

OSWorld is a multimodal benchmark that tests an agent’s ability to complete real tasks in a desktop operating system — navigating GUIs, clicking, typing and using applications the way a person would. Release 2.1 is 108 long-horizon tasks on a live Ubuntu virtual machine, graded against weighted checkpoints. The partial score is the mean per-task credit across checkpoints; the strict pass rate counts only tasks where every checkpoint is satisfied.

Model scores

  • Fable 5.180.7% (partial) / 42.8% (strict)
  • Fable 5—
  • Opus 5.581.8% (partial) / 48.7% (strict)
  • Opus 574.0% (partial) / 37.2% (strict)
  • Opus 4.8—
  • Sonnet 5.580.1% (partial) / 43.5% (strict)
  • Sonnet 557.0% (partial) / 25.6% (strict)
  • GPT-6 Astra—
  • GPT-6.1 Sol—
  • GPT-5.6 Sol—
  • Grok 4.7—
  • GPT-5.5—
  • Composer 2.5—
  • Opus 4.7—
  • Gemini 3.1 Pro—
  • Mythos Preview—

Official source: OSWorld project

Related reading