新基准测试对话智能体在双人协作场景下的表现。
$τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- 构建双向控制环境,模拟用户与智能体共同操作
- 智能体在双人模式下性能显著下降,凸显引导难题
- 适合研究协作式对话系统与人机协同的学者
现有对话智能体评估基准多基于单控制环境,仅智能体可操作工具,用户仅为信息提供者,与真实场景如技术支持不符。为此,我们提出 $τ^2$-bench,包含四项关键贡献:1)建模为分布式部分可观马尔可夫决策过程(Dec-POMDP)的电信双控制领域,智能体与用户共享动态环境并使用工具,考验双方协调与沟通能力;2)组合式任务生成器,从原子组件程序化生成多样且可验证的任务,确保领域覆盖与可控复杂度;3)紧密耦合环境的可靠用户模拟器,其行为受工具与可观测状态约束,提升仿真真实性;4)通过多重消融分析,精细区分推理、沟通与协调错误。实验表明,当智能体从无用户环境切换至双控制时性能大幅下降,凸显引导用户之挑战。$τ^2$-bench 为需兼具有效推理与用户引导能力的智能体提供可控测试平台。
原文摘要 · Abstract (English)
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $τ^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $τ^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。