arXiv:2412.15701cs.AIcs.CL2024-12被引 71

构建人机协作框架,实测协作效果优于纯自主智能体。

Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration

  • 设计非轮流对话的双向交互机制,支持人机协同任务
  • 真实用户测试中协作模型胜率最高达86%(旅行规划)
  • 适合研究人机协作或开发交互式智能应用者使用

尽管大语言模型推动了智能体自动化任务的发展,但许多场景仍需人机协作,以利用人类的隐含偏好、领域专长或控制需求。为此,我们提出 Collaborative Gym(Co-Gym),一个开放框架,用于开发和评估与人类进行双向沟通并交互于任务环境中的协作智能体。该框架通过灵活的非轮流交互范式,支持新任务环境的构建与人机协调,并配备评估套件,用于衡量协作结果与过程。框架提供模拟环境(带可靠用户模拟器)和真实环境(交互式网页应用)。在三个代表性任务——制定旅行计划、撰写相关工作章节、分析表格数据——上的基准实验表明:最佳协作智能体在任务表现上持续优于完全自主模型,真实用户评估下胜率分别为86%(旅行规划)、74%(表格分析)和66%(相关工作)。尽管如此,评估揭示当前语言模型与智能体仍存在局限,真实环境下沟通失败率达65%,情境感知失败率达40%。Co-Gym 以宽松的 MIT 许可证发布,支持新增任务环境,可用于开发协作智能体应用,其评估套件亦可促进协作智能体的优化。

原文摘要 · Abstract (English)

While the advancement of large language models has spurred the development of AI agents to automate tasks, numerous use cases inherently require agents to collaborate with humans due to humans' latent preferences, domain expertise, or the need for control. To facilitate the study of human-agent collaboration, we introduce Collaborative Gym (Co-Gym), an open framework for developing and evaluating collaborative agents that engage in bidirectional communication with humans while interacting with task environments. We describe how the framework enables the implementation of new task environments and coordination between humans and agents through a flexible, non-turn-taking interaction paradigm, along with an evaluation suite that assesses both collaboration outcomes and processes. Our framework provides both a simulated condition with a reliable user simulator and a real-world condition with an interactive web application. Initial benchmark experiments across three representative tasks -- creating travel plans, writing related work sections, and analyzing tabular data -- demonstrate the benefits of human-agent collaboration: The best-performing collaborative agents consistently outperform their fully autonomous counterparts in task performance, achieving win rates of 86% in Travel Planning, 74% in Tabular Analysis, and 66% in Related Work when evaluated by real users. Despite these improvements, our evaluation reveals persistent limitations in current language models and agents, with communication and situational awareness failures observed in 65% and 40% of cases in the real condition, respectively. Released under the permissive MIT license, Co-Gym supports the addition of new task environments and can be used to develop collaborative agent applications, while its evaluation suite enables assessment and improvement of collaborative agents.

人机协作智能体评估双向交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。