评测大模型在真实经济任务中的表现,发现顶尖模型正逼近专家水平。
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- 基于美国劳工部44个职业的典型工作构建评估基准
- 当前最佳模型在交付质量上接近行业专家,且效率更高
- 适合关注大模型真实应用潜力的研究者与从业者
我们提出GDPval,一个评估人工智能模型在真实世界经济价值任务中能力的基准。该基准覆盖美国劳工统计局44个职业、占美国GDP前九大行业的主要工作活动。任务由平均14年经验的行业专业人士设计。我们发现,前沿模型在GDPval上的表现随时间大致呈线性提升,当前最优模型已接近行业专家的交付质量。通过人机协作,前沿模型可比纯人工更快更低成本完成任务。此外,增加推理深度、任务上下文和辅助支持能显著提升模型表现。我们开源了220个高质量任务的金标准数据集,并提供公开自动化评分服务(evals.openai.com),以推动对真实世界模型能力的理解。
原文摘要 · Abstract (English)
We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。