研究多智能体谈判中动态对齐失败问题,发现协作沟通比想象中更难。
Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation

- 设计多轮谈判游戏,测试智能体在互动中建立共同理解的能力。
- 实验显示,即使单个智能体能算出最优解,成对智能体也常无法达成。
- 揭示四类沟通失效模式,强调协作规划与承诺机制的关键作用。
对齐是建立双方共同信念以实现沟通目标的协作过程。静态对齐将语言映射到共享上下文,而动态对齐需在多轮交互中协商意义。当前多智能体大模型基准主要关注静态、一次性任务,忽视了智能体通过互动修复对齐崩溃的能力。我们引入一个迭代式多轮谈判游戏,两个智能体分配共享资源用于私人项目,并存在可验证的联合最优结果。尽管个体智能体可在孤立状态下识别帕累托最优分配,但智能体对在不同模型下始终无法达成该结果。我们识别出四种失败模式:(1)共享交互历史丢失,(2)固守早期提议,(3)默认均分而非追求收益最大化协调,(4)跨轮次指称绑定错误。基线实验表明,协调差距不能仅由个体推理限制或信息交换不足解释,瓶颈在于动态对齐:联合计划形成、承诺与执行。
原文摘要 · Abstract (English)
Grounding is the collaborative process of establishing mutual belief sufficient for a communicative goal. While static grounding maps language to a shared context, dynamic grounding requires agents to negotiate meaning across turns. Current multi-agent Large Language Model (LLM) benchmarks largely emphasize static, one-shot tasks, overlooking whether agents can repair grounding breakdowns through interaction. We introduce an iterated multi-turn negotiation game where two agents allocate shared resources to private projects with verifiable jointly optimal outcomes. Although individual agents can identify Pareto-optimal allocations in isolation, agent dyads consistently fail to reach them across models. We identify four failure modes: (1) loss of shared interaction history, (2) stubborn anchoring to early proposals, (3) defaulting to equal splits over reward-maximizing coordination, and (4) referential binding errors across turns. Our baselines show that the coordination gap is not explained by individual reasoning limits or insufficient information exchange alone. Instead, the bottleneck lies in dynamic grounding: joint plan formation, commitment, and execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。