arXiv:2510.25744cs.CLcs.AI2025-10被引 2

让智能体学会与人协作,而非只求一次完成任务。

Completion $\neq$ Collaboration: Scaling Collaborative Effort with Agents

  • 用多轮交互评估智能体协作能力,而非单次任务完成。
  • 顶尖智能体在真实场景中表现不佳,因缺乏持续引导能力。
  • 适合研究人机协同、交互式AI的开发者与设计者。

当前智能体评估仍聚焦于单次任务完成,未能反映现实问题中目标常不明确且动态演化的迭代协作本质。本文主张从构建任务完成型智能体转向发展协作型智能体,评估不仅看最终输出质量,更关注其在整个问题解决过程中对人类努力的促进与增强。为此,我们提出‘协作努力扩展’框架,量化智能体效用随用户参与度提升的变化。通过案例研究与模拟评估,发现现有先进智能体在多轮真实场景中普遍表现不足,揭示了智能体设计中的关键缺失:维持用户参与并搭建理解的能力。该框架为诊断智能体行为、指导开发更高效交互提供了新视角。

原文摘要 · Abstract (English)

Current evaluations of agents remain centered around one-shot task completion, failing to account for the inherently iterative and collaborative nature of many real-world problems, where human goals are often underspecified and evolve. We argue for a shift from building and assessing task completion agents to developing collaborative agents, assessed not only by the quality of their final outputs but by how well they engage with and enhance human effort throughout the problem-solving process. To support this shift, we introduce collaborative effort scaling, a framework that captures how an agent's utility grows with increasing user involvement. Through case studies and simulated evaluations, we show that state-of-the-art agents often underperform in multi-turn, real-world scenarios, revealing a missing ingredient in agent design: the ability to sustain engagement and scaffold user understanding. Collaborative effort scaling offers a lens for diagnosing agent behavior and guiding development toward more effective interactions.

人机协作多轮交互智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。