Paper 2510.04374
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Published
- Oct 2025
- Research lab
- OpenAI
- Citations
- 108
- GitHub
- Not linked
01 In brief
Summary
This paper introduces GDPval, a benchmark for evaluating AI models on real-world, economically valuable tasks.
It covers 44 occupations across the top 9 U.S.
GDP sectors, with tasks created by industry experts averaging 14 years of experience.
The benchmark includes 1,320 tasks in the full set and a 220-task gold subset, graded via human expert pairwise comparisons.
Results show frontier model performance improving roughly linearly over time, with the best model (Claude Opus 4.1) achieving 47.6% wins or ties against human experts.
Models show potential for cost and time savings when paired with human oversight.
Increased reasoning effort, task context, and scaffolding improve performance.
The authors open-source the gold subset and provide an automated grading service at evals.openai.com.
02 From the paper
Abstract
We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.