arXiv:2605.23262cs.AI2026-05被引 1

为知识型工作设计可比较的评估基准,明确四个关键维度。

Designing Benchmarks for Knowledge Work

论文配图:Designing Benchmarks for Knowledge Work
图 1 · 摘自论文原文
  • 提出四维基准框架:任务活动、测试环境、交付成果、评估结果。
  • 基于O*NET构建18项工作活动清单,验证语义一致性和可解释性。
  • 适用于不同职业的智能体评估,尤其适合研究自动化办公系统。

AI智能体正从单点问答转向通过工具、软件环境和多步骤流程完成工作任务。这类任务多属知识工作,涉及信息解读、专业内容生成与沟通。现有基准通常仅描述任务、环境和指标,隐含四个未明示问题:任务中哪些环节被体现?在何种条件下测试?系统应产出何种工作成果?评估又聚焦成果的哪部分?本文提出以工作为中心的基准表示法,包含四个字段:被代表的活动、测试场景、所需工作产品、评估结果。该框架使设计选择透明可比。为支持跨职业的活动级报告,我们基于O*NET任务语句提炼出18项工作活动,并验证其语义一致性、算法敏感性、在ESCO本体中的可读性及人工可理解性。将该框架应用于GDPval、OfficeQA Pro和APEX-SWE三个基准,案例分析表明:职业交付物、准确回答与可执行状态变化,可在同一框架下捕捉工作不同侧面。

原文摘要 · Abstract (English)

AI agents are moving quickly from answering isolated questions toward completing work through tools, software environments, and multi-step workflows. Much of what these systems are now asked to do is knowledge work, where information and expertise are interpreted, produced, and communicated as part of completing work. Benchmarks for this setting are usually described only by their tasks, environments, and metrics, leaving four questions implicit: what part of the work is represented, under what conditions it is tested, what work product the system is expected to leave, and what part of that product the benchmark actually evaluates. We introduce a work-centered benchmark representation with four fields: represented activity, tested setting, required work product, and evaluated result. The representation makes these choices explicit and comparable across benchmark designs. To support activity-level reporting across occupations, we derive an aim-dependent inventory of 18 work activities from O*NET task statements and report evidence on semantic coherence, algorithm sensitivity, external ontology legibility in ESCO, and human interpretability. We apply the representation to GDPval, OfficeQA Pro, and APEX-SWE. The case analyses illustrate how occupational deliverables, grounded answers, and executable state changes capture different parts of work within the same representation.

智能体评估知识工作基准设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。