让多个智能体在无通信情况下学会协同完成多解任务。
MIMIC-D: Multi-modal Imitation for MultI-agent Coordination with Decentralized Diffusion Policies
- 用扩散模型联合训练各智能体策略,仅依赖局部信息
- 实现在仿真与真实硬件中多种任务的多模态协作
- 适合无法直接通信的机器人或人机协作场景
随着机器人融入社会,其在多模态任务(存在多种有效解)中与他人协同的能力至关重要。传统模仿学习在面对多模态示范时往往平均化或退化到单一模式,难以实现有效协作。受单智能体扩散模型捕捉复杂轨迹分布能力启发,本文提出MIMIC-D框架,通过扩散模型实现多智能体系统的多模态协同行为。现有方法通常依赖中心化规划或显式通信,而真实场景中智能体常需独立运行或与无法直接通信的个体(如人类)协作。为此,我们设计了一种联合训练、去中心化执行的范式:所有智能体仅用本地信息联合训练,实现隐式协调。在仿真和硬件实验中,该方法在多种任务与环境中展现出稳健的多模态协作性能,优于当前最优基线。
原文摘要 · Abstract (English)
As robots become more integrated in society, their ability to coordinate with other robots and humans on multi-modal tasks (those with multiple valid solutions) is crucial. Such behaviors can be learned from expert demonstrations via imitation learning (IL), but when expert demonstrations are multi-modal, standard IL approaches usually average across modes or collapse to a single mode, preventing effective coordination. Being inspired by diffusion models' ability to capture complex multi-modal trajectory distributions in single-agent settings, we develop a diffusion-based framework for coordinated multi-modal behavior in multi-agent systems. However, existing multi-agent diffusion approaches typically require a centralized planner or explicit communication among agents. This assumption can fail in real-world scenarios where robots must operate independently or with agents like humans that they cannot directly communicate with. Therefore, we propose MIMIC-D, a joint training with decentralized execution paradigm for multi-modal multi-agent IL via diffusion. We jointly train all agents' policies with only local information to achieve implicit coordination. In simulation and hardware experiments, our method exhibits robust multi-modal coordination behavior in various tasks and environments, improving upon state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。