arXiv:2502.18438cs.MAcs.AI2025-02

让智能体团队具备心理理论能力,动态生成协作计划。

ToMCAT: Theory-of-Mind for Cooperative Agents in Teams via Multiagent Diffusion Policies

  • 用元学习推断队友目标与行为,结合扩散模型生成协同轨迹。
  • 动态重规划显著降低资源消耗,同时保持团队性能不降。
  • 适合无先验信息的动态协作场景,尤其在复杂环境中表现优异。

本文提出ToMCAT(团队协作智能体的心理理论框架),通过融合元学习机制与多智能体去噪扩散模型,实现基于心理理论(ToM)的协同轨迹生成。该框架能推断队友潜在目标与未来行为,并生成与自身目标及队友特征相匹配的联合计划。我们构建了在线重规划系统,当发现计划与当前环境状态偏离时,会动态从扩散模型中采样新轨迹。在模拟烹饪任务中的实验表明,动态重规划可有效减少资源使用,且不牺牲团队整体表现。此外,结合当前观测与心理理论推理,对生成适应性强的团队计划至关重要,尤其在缺乏队友先验信息的情况下。

原文摘要 · Abstract (English)

In this paper we present ToMCAT (Theory-of-Mind for Cooperative Agents in Teams), a new framework for generating ToM-conditioned trajectories. It combines a meta-learning mechanism, that performs ToM reasoning over teammates' underlying goals and future behavior, with a multiagent denoising-diffusion model, that generates plans for an agent and its teammates conditioned on both the agent's goals and its teammates' characteristics, as computed via ToM. We implemented an online planning system that dynamically samples new trajectories (replans) from the diffusion model whenever it detects a divergence between a previously generated plan and the current state of the world. We conducted several experiments using ToMCAT in a simulated cooking domain. Our results highlight the importance of the dynamic replanning mechanism in reducing the usage of resources without sacrificing team performance. We also show that recent observations about the world and teammates' behavior collected by an agent over the course of an episode combined with ToM inferences are crucial to generate team-aware plans for dynamic adaptation to teammates, especially when no prior information is provided about them.

多智能体心理理论扩散模型动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。