Terminal-Bench 4.0
Terminal-Bench evaluates models on real-world tasks in a terminal and command-line environment — installing dependencies, debugging, running builds and orchestrating tools — where each step depends on the result of the last. Release 4.0 is the current task set, and the project publishes a single leaderboard on which every vendor runs its own agent harness. Scores are resolution rates with a 95% confidence interval.
Official source: Terminal-Bench 4.0 leaderboard (tbench.ai)