Paper 2606.05405
Agents' Last Exam
- Published
- Jun 2026
- Research lab
- Independent
- Citations
- 5
- GitHub
- 936 stars
01 In brief
Summary
Agents' Last Exam (ALE) is a benchmark introduced by UC Berkeley and collaborators to evaluate AI agents on long-horizon, economically valuable, real-world professional tasks with verifiable outcomes.
Developed with 250+ industry experts, ALE covers 55 subfields across 13 industry clusters, grounded in the O*NET/SOC 2018 occupational taxonomy, and includes 1,490 task instances (960 expert submissions, 530 commissioned).
Tasks are sourced from real projects completed by practitioners and undergo a five-gate quality control pipeline.
Evaluation targets Generalist Computer-Use Agents (GCUA) that combine GUI, CLI, and tool use.
Results show the benchmark is far from saturated: the strongest configuration (Codex with GPT-5.5) achieves only 24.0% overall pass rate, with near-zero pass rates on the hardest tier.
ALE is designed as a living benchmark with a public/private split (150 public, 1,017 private, 323 unverified) to mitigate contamination, and aims to close the gap between benchmark success and GDP-relevant impact.
The paper details the task construction pipeline, evaluation architecture, and extensive experimental results across multiple harnesses and models, including a failure analysis showing that domain knowledge gaps and wrong strategies dominate failures.
ALE-CLI, a Linux-only subset, is also introduced for comparison with CLI-only agents.
The benchmark is positioned as the first to cover all 55 SOC/O*NET industries with deterministic, rubric-based automated verification, unlike prior benchmarks that rely on human grading or cover fewer domains.
The authors release ALE as an instrument for measuring AI's readiness for real professional work, where saturation would signal economic transformation.
Funding was provided by the Tianqiao & Chrissy Chen…
02 From the paper
Abstract
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.