arXiv:2605.26329cs.AI2026-05被引 4

新基准评估AI代理在真实职业任务中的协作能力,强调人类需求而非经济价值。

JobBench: Aligning Agent Work With Human Will

论文配图:JobBench: Aligning Agent Work With Human Will
图 1 · 摘自论文原文
  • 以专家认定的高优先级任务为评估核心,模拟真实工作场景中的信息杂乱环境。
  • 36个模型中最强表现仅达45.9%,反映当前代理在复杂任务中仍严重不足。
  • 适合关注AI增强人类而非替代的科研者与从业者参考。

现有职业类AI代理评测主要基于经济价值,暗示取代人类。我们提出JobBench,评估AI代理在专家认定需委托的130项任务(覆盖35个职业)上的表现,聚焦人类实际需求而非GDP价值。每项任务以包含异构参考文件的工作空间形式呈现,要求代理处理真实职场中混乱的信息流。评价采用基于事实的链式评分标准,平均每任务35.6个二元评判项。评估36个模型,最优结果为Claude Opus 4.7在Claude Code下达到45.9%。我们希望JobBench引导社区从取代转向增强:构建真正实现人类意愿委托的智能体,而非仅追求经济价值最高的任务。

原文摘要 · Abstract (English)

Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task. We evaluate 36 models; the strongest, Claude Opus~4.7 under Claude Code, reaches only 45.9 %. We hope JobBench shifts the community's target labour-market effect from replacement to enhancement: building agents that do what humans actually want delegated, not only what is most economically valuable.

AI代理职业任务人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。