让大模型代理自动匹配对方思维深度,提升协作效率。
Adaptive Theory of Mind for LLM-based Multi-Agent Coordination
- 根据交互历史预测伙伴的思维层级,动态调整自身推理深度。
- 在4个协作任务中,自适应思维模型显著优于固定思维层级方案。
- 适用于各类智能体,尤其适合需要精准协作的复杂场景。
心智理论(ToM)指推断他人心理状态的能力,高阶ToM则涉及理解他人也具备自己的心智理论。赋予大语言模型驱动的智能体以ToM能力,长期被认为可提升多智能体协作表现。然而我们发现,若各智能体的ToM层级不匹配——即对他人思维深度的推理存在偏差——将导致推理不足或过度,从而损害协作效果。为此,我们设计了自适应心智理论(A-ToM)智能体,能与合作方对齐思维层级。基于前期互动,该智能体估计伙伴可能的ToM层级,并利用此估计预测其行为,实现行为协调。我们在四个多智能体协作任务上进行实证评估:重复矩阵博弈、两个网格导航任务和Overcooked任务。结果验证了ToM对齐的重要性,并证明了A-ToM的有效性。此外,我们讨论了A-ToM在非LLM智能体中的泛化潜力,以及何种情况下ToM对齐的重要性会降低。
原文摘要 · Abstract (English)
Theory of Mind (ToM) refers to the ability to reason about others' mental states, and higher-order ToM involves considering that others also possess their own ToM. Equipping large language model (LLM)-driven agents with ToM has long been considered to improve their coordination in multiagent collaborative tasks. However, we find that misaligned ToM orders-mismatches in the depth of ToM reasoning between agents-can lead to insufficient or excessive reasoning about others, thereby impairing their coordination. To address this issue, we design an adaptive ToM (A-ToM) agent, which can align in ToM orders with its partner. Based on prior interactions, the agent estimates the partner's likely ToM order and leverages this estimation to predict the partner's action, thereby facilitating behavioral coordination. We conduct empirical evaluations on four multi-agent coordination tasks: a repeated matrix game, two grid navigation tasks and an Overcooked task. The results validate our findings on ToM alignment and demonstrate the effectiveness of our A-ToM agent. Furthermore, we discuss the generalizability of our A-ToM to non-LLM-based agents, as well as what would diminish the importance of ToM alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。