arXiv:2509.25721econ.GNcs.AI2025-09被引 10

评测顶尖模型在四大职业中的实际工作能力,发现仍存显著短板。

The AI Productivity Index (APEX)

  • 扩展基准测试至400个真实任务案例,覆盖四大专业岗位。
  • GPT5(高思考模式)得分67.0%,仍难胜任多数专业任务。
  • 开源100个非测试案例与评估工具,支持后续研究。

我们提出扩展版人工智能生产力指数(APEX-v1-extended),用于评估前沿模型在投资银行分析师、管理咨询顾问、大型律所律师及初级医生(MD)四类职业中完成经济价值任务的能力。本技术报告详细说明了APEX-v1的改进:将保留测试集从每岗位50例增至100例(总计400例),并更新评分方法。新排行榜显示,GPT5(高思考模式)以67.0%的得分位居榜首。结果表明,当前前沿模型在执行典型专业任务时仍存在明显局限。为促进后续研究,我们开放源代码,提供每岗位25个非基准示例案例(共100例)及评估框架。

原文摘要 · Abstract (English)

We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

AI评估生产力职业能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。