arXiv:2511.02762cs.LGcs.MA2025-11被引 1

用单人经验训练团队智能体,提升协作效率。

From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos

  • 从单人示范中预训练共享策略,再通过融合机制适配团队协作。
  • 在多种任务上显著提升训练效率与性能,超越基线模型。
  • 适合缺乏多智能体数据的场景,如协同编程、救援任务等。

在多智能体强化学习中从零训练团队智能体效率极低,如同让新手直接合奏交响乐。现有方法虽缓解此问题,但仍依赖昂贵的多智能体数据。而单人经验在协作编程、家庭合作、搜救等场景中更易获取。为此,我们提出SoCo框架,将单人知识迁移至协作学习:先从单人示范中预训练共享策略,再通过类似MoE的门控选择器与动作编辑器,在多智能体训练中实现策略融合。在多种协作任务上的实验表明,SoCo显著提升骨干算法的训练效率与性能。结果证明,单人示范可作为多智能体数据的有效补充,使协作学习更具可扩展性与实用性。

原文摘要 · Abstract (English)

Training a team of agents from scratch in multi-agent reinforcement learning (MARL) is highly inefficient, much like asking beginners to play a symphony together without first practicing solo. Existing methods, such as offline or transferable MARL, can ease this burden, but they still rely on costly multi-agent data, which often becomes the bottleneck. In contrast, solo experiences are far easier to obtain in many important scenarios, e.g., collaborative coding, household cooperation, and search-and-rescue. To unlock their potential, we propose Solo-to-Collaborative RL (SoCo), a framework that transfers solo knowledge into cooperative learning. SoCo first pretrains a shared solo policy from solo demonstrations, then adapts it for cooperation during multi-agent training through a policy fusion mechanism that combines an MoE-like gating selector and an action editor. Experiments across diverse cooperative tasks show that SoCo significantly boosts the training efficiency and performance of backbone algorithms. These results demonstrate that solo demonstrations provide a scalable and effective complement to multi-agent data, making cooperative learning more practical and broadly applicable.

多智能体强化学习知识迁移协作训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。