通过交互式教学提升多智能体强化学习的协作效率
Interactive Distillation for Cooperative Multi-Agent Reinforcement Learning
- 采用分层教学框架,让教师同时利用自身与学生经验进行优化
- 在复杂任务中成功率提升60%至165%,显著优于现有方法
- 适合需要高效协作的多智能体系统,如资源分配与战术对抗
知识蒸馏(KD)有望通过集中式教师指导分布式学生来加速多智能体强化学习(MARL),但面临三大瓶颈:(1)在复杂环境中合成高性能教学策略困难;(2)当教师需在分布外(OOD)状态推理时表现不佳;(3)学生与教师观测空间不匹配。为此,我们提出HINT(Hierarchical INteractive Teacher-based transfer),一种面向集中训练、分散执行场景的新型KD框架。借助分层强化学习,HINT构建可扩展且高性能的教师策略。其核心创新——伪离线策略学习,使教师能融合自身与学生的经验进行更新,从而增强对分布外状态的适应能力。HINT还引入基于性能的过滤机制,仅保留与结果相关的信息,缓解观测空间差异。我们在多个挑战性协作任务上进行了评估,包括FireCommander(资源分配)和MARINE(战术对抗)。实验表明,相较于基线方法,HINT在各项基准测试中成功率达60%至165%的提升。
原文摘要 · Abstract (English)
Knowledge distillation (KD) has the potential to accelerate MARL by employing a centralized teacher for decentralized students but faces key bottlenecks. Specifically, there are (1) challenges in synthesizing high-performing teaching policies in complex domains, (2) difficulties when teachers must reason in out-of-distribution (OOD) states, and (3) mismatches between the decentralized students' and the centralized teacher's observation spaces. To address these limitations, we propose HINT (Hierarchical INteractive Teacher-based transfer), a novel KD framework for MARL in a centralized training, decentralized execution setup. By leveraging hierarchical RL, HINT provides a scalable, high-performing teacher. Our key innovation, pseudo off-policy RL, enables the teacher policy to be updated using both teacher and student experience, thereby improving OOD adaptation. HINT also applies performance-based filtering to retain only outcome-relevant guidance, reducing observation mismatches. We evaluate HINT on challenging cooperative domains (e.g., FireCommander for resource allocation, MARINE for tactical combat). Across these benchmarks, HINT outperforms baselines, achieving improvements of 60% to 165% in success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。