用单智能体数据生成多智能体协同行为,无需联合演示。
Coordinated Diffusion: Generating Multi-Agent Behavior Without Multi-Agent Demonstrations

- 通过成本函数引导独立训练的单智能体扩散模型协同
- 仅用单智能体数据实现鲁棒的双臂协同操作,效率高于基线
- 适用于黑箱成本函数,适合机器人协同控制场景
基于生成模型的模仿学习在建模复杂单智能体行为方面表现优异。然而,由于联合状态-动作空间随智能体数量呈指数增长,收集足够多的协同多智能体示范数据极为昂贵,制约了多智能体系统的模仿学习。本文提出Coordinated Diffusion(CoDi),一种通过用户定义的多智能体成本函数耦合独立训练的单智能体扩散策略的框架,无需任何协同示范数据。我们推导出一种新的基于扩散的采样方案,其中扩散得分函数分解为独立的单智能体预训练基础策略加上由成本驱动的引导项,该引导项协调各基础策略形成一致的多智能体行为。该引导项可无梯度估计,使CoDi适用于黑箱、不可微分的成本函数且无需额外训练。理论上和实证上分析了该组合能忠实逼近目标多智能体行为的条件。结果表明:单智能体示范数据需覆盖目标多智能体行为的支持集,而成本函数则需从基础策略的乘积中促进期望行为。仿真与硬件实验在双臂操作任务中验证了CoDi能从单智能体数据中发现稳健的协同行为,数据效率优于多智能体基线,并凸显了联合引导、基础策略支持与成本设计的重要性。
原文摘要 · Abstract (English)
Imitation learning powered by generative models has proven effective for modeling complex single-agent behaviors. However, teaching multi-agent systems, like multiple arms or vehicles, to coordinate through imitation learning is hindered by a fundamental data bottleneck: as the joint state-action space grows exponentially with the number of agents, collecting a sufficient amount of coordinated multi-agent demonstrations becomes extremely costly. In this work, we ask: how can we leverage single-agent demonstration data to learn multi-agent policies? We present Coordinated Diffusion (CoDi), a framework that couples independently trained single-agent diffusion policies through a user-defined multi-agent cost function, without requiring any coordinated demonstrations. We derive a new diffusion-based sampling scheme wherein the diffusion score function decomposes into independent, single-agent pre-trained base policies plus a cost-driven guidance term that coordinates these base policies into cohesive multi-agent behavior. We show that this guidance term can be estimated in a gradient-free manner, making CoDi applicable to black-box, non-differentiable cost functions without additional training. Theoretically and empirically, we analyze the conditions under which this composition can faithfully approximate a target multi-agent behavior. We find a complementary role for demonstration data versus the cost function: single-agent demonstrations must cover the support of the desired multi-agent behavior, while the cost function must promote desired behavior from this product of single-agent policies. Our results in simulation and hardware experiments of a two-arm manipulation task show that CoDi discovers robust coordinated behavior from single-agent data, is more data-efficient than multi-agent baselines, and highlights the importance of joint guidance, base policy support, and cost design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。