OpenAI tests models against specialists from 44 professions

OpenAI introduced new benchmark GDPval, which tests its AI models’ performance compared to professionals from various industries. And is an attempt to understand how close OpenAI systems are to surpassing humans in economically significant work.

The benchmark is based on 9 industries making largest contribution to US gross domestic product. GDPval tests AI model performance across 44 professions in these industries, from programmers to nurses and journalists. Experienced professionals compared AI-generated reports with works of other specialists.

GPT-5 high was rated better than or equal to industry experts in 46.6% of cases. Claude Opus 4.1 from Anthropic was rated better than or equal to industry experts in 49% of tasks. Although OpenAI claims Claude showed such high results due to tendency to create attractive graphics.

I think such high model scores might be inflated due to test limitations. And don’t reflect real performance. The new benchmark itself could create false expectations about AI capabilities in real work conditions.