arXiv:2604.10866cs.CL2026-04被引 4

用语言模拟器构建100个职业任务基准,评测AI在真实场景中的表现。

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

论文配图:OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
图 1 · 摘自论文原文
  • 通过大模型生成工具响应,模拟专业环境进行评估。
  • 隐性故障比显性错误更难,需自主发现数据缺失。
  • 大模型和高推理能力显著提升表现,适合行业应用研究者。

AI代理被期望在数百个职业领域(从急诊分诊到核反应堆安全监控再到海关进口处理)完成专业工作,但现有基准仅能评估少数有公开环境的领域。我们提出OccuBench,涵盖10个行业类别下65个专业领域的100个真实职业任务场景,依托语言环境模拟器(LES)通过大模型驱动的工具响应生成来模拟领域特定环境。多智能体合成流程自动生成保证可解性、难度校准且文档驱动多样性的评估实例。OccuBench从两个互补维度评估代理:跨职业领域的任务完成率,以及在受控故障注入下的环境鲁棒性(显式错误、隐式数据退化、混合故障)。我们在8个模型家族的15个前沿模型上进行评估,发现:(1) 没有单一模型在所有行业中占优,每种模型均有独特的职业能力画像;(2) 隐性故障(数据截断、字段缺失)比显式错误(超时、500错误)和混合故障更难,因其缺乏明显错误信号,需代理自主检测数据退化;(3) 更大模型、更新世代和更高推理投入均持续提升性能,GPT-5.2在推理努力从最低到最高时提升27.5分;(4) 强代理未必是强环境模拟器,模拟器质量对基于LES的评估可靠性至关重要。OccuBench提供了首个系统性的跨行业职业任务评估。

原文摘要 · Abstract (English)

AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety monitoring to customs import processing), yet existing benchmarks can only evaluate agents in the few domains where public environments exist. We introduce OccuBench, a benchmark covering 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, enabled by Language Environment Simulators (LESs) that simulate domain-specific environments through LLM-driven tool response generation. Our multi-agent synthesis pipeline automatically produces evaluation instances with guaranteed solvability, calibrated difficulty, and document-grounded diversity. OccuBench evaluates agents along two complementary dimensions: task completion across professional domains and environmental robustness under controlled fault injection (explicit errors, implicit data degradation, and mixed faults). We evaluate 15 frontier models across 8 model families and find that: (1) no single model dominates all industries, as each has a distinct occupational capability profile; (2) implicit faults (truncated data, missing fields) are harder than both explicit errors (timeouts, 500s) and mixed faults, because they lack overt error signals and require the agent to independently detect data degradation; (3) larger models, newer generations, and higher reasoning effort consistently improve performance. GPT-5.2 improves by 27.5 points from minimal to maximum reasoning effort; and (4) strong agents are not necessarily strong environment simulators. Simulator quality is critical for LES-based evaluation reliability. OccuBench provides the first systematic cross-industry evaluation of AI agents on professional occupational tasks.

AI代理职业任务评估基准语言模拟器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。