arXiv:2606.25389cs.AI2026-06被引 1

通过技能拆分与复用,让多智能体持续学习协作技能

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

论文配图:Offline Multi-agent Continual Cooperation via Skill Partition and Reuse
图 1 · 摘自论文原文
  • 从离线数据中自动提取可复用的协作技能
  • 在多个任务流中持续扩展技能库,减少干扰和遗忘
  • 适合需要长期协作学习的开放环境多智能体系统

从多智能体离线数据中提取技能,可通过共享任务无关的协调技能提升学习效率。在任务顺序发生且技能空间呈指数增长的场景下,依赖启发式设计且固定大小的技能库的方法难以应对分布偏移与干扰问题,导致灾难性遗忘和可塑性损失。为此,我们提出COMAD框架,基于技能拆分与复用,实现持续的离线多智能体技能发现。首先利用自编码器从混合多智能体行为数据中提取协调知识,并转化为可复用的协调技能;随后构建基于多头架构的技能增强策略学习目标,通过基于密度的可复用性估计器显式引导优势函数。理论分析表明,该方法逼近持续技能发现问题的最优解。在多种多智能体强化学习基准测试中,COMAD持续扩展技能库以缓解干扰,在任务流上实现了优于多个基线的前向与后向迁移性能。

原文摘要 · Abstract (English)

Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where tasks occur sequentially and the space of skills grows exponentially, existing approaches that rely on heuristically designed and fixed-sized skill libraries struggle to resolve the problem of distributional shift and interference, facing catastrophic forgetting and plasticity loss. To address this problem and endow agents with the ability to continually discover and reuse coordination skills in open-environment, we propose COMAD, a principled framework for Continual Offline Multi-agent Skill Discovery via Skill Partition and Reuse. We first discover skills from mixed multi-agent behavior data with an auto-encoder to transform coordination knowledge into reusable coordination skills. Then we construct a skill-augmented policy learning objective with multi-head architectures, explicitly guiding the advantage function with reusable skills identified via a density-based reusability estimator. Theoretical analysis shows our method approximates the optimum of a continual skill discovery problem. Empirical results across diverse MARL benchmarks show that COMAD continually expands its skill library to mitigate interference, achieving superior forward and backward transfer for task streams compared to multiple baselines.

多智能体持续学习技能复用强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。