arXiv:2607.04329cs.AI2026-07

评测大模型与人协作时,不同参与方式对任务效果的影响。

HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

论文配图:HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
图 1 · 摘自论文原文
  • 用图结构建模人与智能体的职责、权限和沟通路径。
  • 实验证明人类参与能提升任务完成率和容错能力。
  • 适合研究人机协同系统或评估交互设计的学者使用。

大型语言模型越来越多地应用于人类作为主动合作者而非被动任务提供者的场景。我们提出HAS-Framework,一种基于图的框架,将人类与大模型驱动的智能体均视为具有明确角色、权限、通信路径和行动权的第一类参与者。在此框架基础上,HAS-Bench 在可配置的人类参与条件下,评估了不同代理层级、交互渠道和人格策略下的人机系统表现。该基准不仅衡量任务结果,还评估过程层面的合作行为,包括澄清质量、反馈利用、控制校准、安全性、主动性及交互成本。在六个领域的实验表明,人类参与可显著提升任务完成率与故障恢复能力,但收益取决于人类输入的时机、方式及执行者。

原文摘要 · Abstract (English)

Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework that represents humans and LLM-powered agents as first-class participants with explicit roles, permissions, communication paths, and action authority. Building on this framework, HAS-Bench evaluates Human-Agent Systems under configurable human participation across agency levels, interaction channels, and persona policies. The benchmark measures both task outcomes and process-level collaboration behavior, including clarification quality, feedback utilization, control calibration, safety, initiative, and interaction cost. Experiments across six domains show that human participation can substantially improve task completion and failure recovery, but the gains depend on when, how, and by whom human input is exercised.

人机协同评估基准大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。