让智能体零样本推断他人意图,实现无需通信的协作
MetaMind: General and Cognitive World Models in Multi-Agent Systems by Meta-Theory of Mind
- 通过自反性双向推理,每个智能体可自我学习认知能力
- 在多智能体任务中实现少样本泛化,性能超越基线方法
- 适合研究自主协作系统或具身智能的学者参考
多智能体系统的世界模型面临理解相互依赖的智能体动态、预测交互轨迹并进行长时程规划的挑战,且无需中心化监督或显式通信。本文提出MetaMind,一种基于新型元心智(Meta-ToM)框架的一般性认知世界模型。通过该模型,每个智能体不仅能预测和规划自身信念,还能从自身行为轨迹中逆向推理目标与信念,形成自监督的元认知能力。这种自反性双向推理机制使智能体能通过类比推理,将第一人称认知能力泛化至第三人称视角。因此,在多智能体系统中,每个具备MetaMind的智能体可零样本地从有限可观测行为轨迹中主动推理他者的目标与信念,并在无显式通信下适应涌现的集体意图。在多种多智能体任务上的扩展仿真结果表明,MetaMind在任务性能上表现优异,且在少样本多智能体泛化方面优于基线方法。
原文摘要 · Abstract (English)
A major challenge for world models in multi-agent systems is to understand interdependent agent dynamics, predict interactive multi-agent trajectories, and plan over long horizons with collective awareness, without centralized supervision or explicit communication. In this paper, MetaMind, a general and cognitive world model for multi-agent systems that leverages a novel meta-theory of mind (Meta-ToM) framework, is proposed. Through MetaMind, each agent learns not only to predict and plan over its own beliefs, but also to inversely reason goals and beliefs from its own behavior trajectories. This self-reflective, bidirectional inference loop enables each agent to learn a metacognitive ability in a self-supervised manner. Then, MetaMind is shown to generalize the metacognitive ability from first-person to third-person through analogical reasoning. Thus, in multi-agent systems, each agent with MetaMind can actively reason about goals and beliefs of other agents from limited, observable behavior trajectories in a zero-shot manner, and then adapt to emergent collective intention without an explicit communication mechanism. Extended simulation results on diverse multi-agent tasks demonstrate that MetaMind can achieve superior task performance and outperform baselines in few-shot multi-agent generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。