arXiv:2604.06691cs.AI2026-04中稿 · IJCNN 2026被引 1

让多智能体强化学习在资源受限设备上高效运行

KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

  • 分两阶段将专家的协作行为迁移到轻量学生模型
  • 性能保留超90%,计算量降低最多28.6倍
  • 适合边缘设备部署,支持异构智能体能力

现实世界中多智能体强化学习(MARL)系统的部署受限于计算资源、内存和推理时间。尽管专家策略性能优异,但依赖高成本决策周期和大规模模型,难以在边缘设备或嵌入式平台使用。知识蒸馏(KD)为资源感知执行提供了可行路径,但现有方法多聚焦于动作模仿,忽视协作结构且假设智能体能力一致。本文提出资源感知的多智能体知识蒸馏框架(KD-MARL),通过两阶段流程将集中式专家的协调行为迁移至轻量化分布式学生智能体。学生策略训练无需评论家,而是依靠蒸馏的优势信号与结构化策略监督,在异构观测条件下保持协作能力。该方法同时传递动作行为与结构化协作模式,支持异构学生架构,使每个智能体的模型容量匹配其观测复杂度,对部分可观测及资源受限场景至关重要。在SMAC和MPE基准上的实验表明,KD-MARL在保持超过90%专家性能的同时,将计算成本降低最高达28.6倍(FLOPs)。该方法实现了专家级协作并经结构化蒸馏得以保留,推动了资源受限平台上MARL的实际应用。

原文摘要 · Abstract (English)

Real world deployment of multi agent reinforcement learning MARL systems is fundamentally constrained by limited compute memory and inference time. While expert policies achieve high performance they rely on costly decision cycles and large scale models that are impractical for edge devices or embedded platforms. Knowledge distillation KD offers a promising path toward resource aware execution but existing KD methods in MARL focus narrowly on action imitation often neglecting coordination structure and assuming uniform agent capabilities. We propose resource aware Knowledge Distillation for Multi Agent Reinforcement Learning KD MARL a two stage framework that transfers coordinated behavior from a centralized expert to lightweight decentralized student agents. The student policies are trained without a critic relying instead on distilled advantage signals and structured policy supervision to preserve coordination under heterogeneous and limited observations. Our approach transfers both action level behavior and structural coordination patterns from expert policies while supporting heterogeneous student architectures allowing each agent model capacity to match its observation complexity which is crucial for efficient execution under partial or limited observability and limited onboard resources. Extensive experiments on SMAC and MPE benchmarks demonstrate that KD MARL achieves high performance retention while substantially reducing computational cost. Across standard multi agent benchmarks KD MARL retains over 90 percent of expert performance while reducing computational cost by up to 28.6 times FLOPs. The proposed approach achieves expert level coordination and preserves it through structured distillation enabling practical MARL deployment across resource constrained onboard platforms.

多智能体知识蒸馏边缘计算资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。