arXiv:2603.01203cs.AI2026-03被引 8

研究发现当前AI代理开发与真实工作不匹配,提出三大评估设计原则。

How Well Does Agent Development Reflect Real-World Work?

  • 将43个基准任务映射到1016种职业,分析其与真实劳动市场的契合度。
  • 发现代理开发过度聚焦编程,而人类工作和经济价值集中于其他领域。
  • 提出覆盖、真实性和细粒度评估三原则,指导更贴近社会需求的基准设计。

AI代理在模拟人类工作方面的发展日益增多,但其评估基准是否反映真实劳动力市场仍不明确。本研究系统分析了代理开发与真实人类工作分布之间的关系,通过将基准任务映射至工作领域和技能,分析43个基准中的72,342项任务,衡量其与美国1,016种职业中的人力就业与资本配置的对齐程度。研究揭示出代理开发高度集中于编程类任务,而人类劳动和经济价值主要集中在其他领域,存在显著偏差。在代理已关注的工作领域内,进一步通过自主性水平测量其实际可用性,为不同工作场景下的交互策略提供实践指导。基于此,提出三个可量化的基准设计原则:覆盖性(coverage)、真实性(realism)和细粒度评估(granular evaluation),以更好捕捉具有社会重要性与技术挑战性的实际工作。

原文摘要 · Abstract (English)

AI agents are increasingly developed and evaluated on benchmarks relevant to human work, yet it remains unclear how representative these benchmarking efforts are of the labor market as a whole. In this work, we systematically study the relationship between agent development efforts and the distribution of real-world human work by mapping benchmark instances to work domains and skills. We first analyze 43 benchmarks and 72,342 tasks, measuring their alignment with human employment and capital allocation across all 1,016 real-world occupations in the U.S. labor market. We reveal substantial mismatches between agent development that tends to be programming-centric, and the categories in which human labor and economic value are concentrated. Within work areas that agents currently target, we further characterize current agent utility by measuring their autonomy levels, providing practical guidance for agent interaction strategies across work scenarios. Building on these findings, we propose three measurable principles for designing benchmarks that better capture socially important and technically challenging forms of work: coverage, realism, and granular evaluation.

AI代理评估基准劳动市场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。