arXiv:2503.21200cs.LGcs.AI2025-03ICLR被引 16

通过分层技能学习,让多智能体在离线数据中学会通用协作能力。

Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent Cooperation

  • 分层设计:同时学习通用技能和任务专属技能
  • 在未见过的任务上实现更优协作表现
  • 适合需要灵活适应多变场景的多智能体系统

从离线多任务数据中学习可泛化的多智能体协作策略,以应对未见任务中智能体数量与目标数变化的挑战,具有重要意义。尽管将多个任务中的共性行为模式抽象为技能以提升策略迁移能力是可行方向,但现有方法面临两大障碍:一是从不同动作序列中提取的通用协作行为缺乏协同的时间知识;二是仅依赖通用技能,无法在每项任务中自适应选择独立的特定知识以实现精细动作执行。为此,我们提出分层分离技能发现(HiSSD)方法,通过分层框架联合学习通用技能与任务专属技能。通用技能捕捉协作的时间规律,支持离线多任务强化学习中的样本内利用;任务专属技能体现各任务先验信息,实现任务引导的精细化动作控制。在多智能体MuJoCo与SMAC基准测试中验证表明,使用HiSSD训练后的策略能有效分配协作行为,在未见任务上取得显著性能提升。

原文摘要 · Abstract (English)

Learning cooperative multi-agent policy from offline multi-task data that can generalize to unseen tasks with varying numbers of agents and targets is an attractive problem in many scenarios. Although aggregating general behavior patterns among multiple tasks as skills to improve policy transfer is a promising approach, two primary challenges hinder the further advancement of skill learning in offline multi-task MARL. Firstly, extracting general cooperative behaviors from various action sequences as common skills lacks bringing cooperative temporal knowledge into them. Secondly, existing works only involve common skills and can not adaptively choose independent knowledge as task-specific skills in each task for fine-grained action execution. To tackle these challenges, we propose Hierarchical and Separate Skill Discovery (HiSSD), a novel approach for generalizable offline multi-task MARL through skill learning. HiSSD leverages a hierarchical framework that jointly learns common and task-specific skills. The common skills learn cooperative temporal knowledge and enable in-sample exploitation for offline multi-task MARL. The task-specific skills represent the priors of each task and achieve a task-guided fine-grained action execution. To verify the advancement of our method, we conduct experiments on multi-agent MuJoCo and SMAC benchmarks. After training the policy using HiSSD on offline multi-task data, the empirical results show that HiSSD assigns effective cooperative behaviors and obtains superior performance in unseen tasks.

多智能体离线学习技能发现泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。