arXiv:2502.00870cs.LGcs.AI2025-02中稿 · AAMAS 2025, includ…被引 8

解决异构联邦强化学习中的知识共享难题,无需公开数据集

FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation

  • 用动作概率分布作为异构智能体间知识传递的媒介
  • 在多个基准任务上显著提升性能,收敛性有理论保障
  • 不依赖复杂公共数据集,适合真实多样的联邦场景

联邦强化学习(FedRL)在保护隐私的同时提升样本效率;然而,现有研究大多假设智能体同质,限制了其在真实场景的应用。本文研究黑箱环境下异构智能体的联邦强化学习,各智能体使用不同的策略网络和训练配置,且不公开内部细节。知识蒸馏(KD)虽是异构模型间知识共享的可行方法,但在联邦强化学习中面临公共数据稀缺与知识表征受限的问题。为此,我们提出联邦异构策略蒸馏(FedHPD),通过动作概率分布作为知识传递媒介,解决异构联邦强化学习问题。我们在标准假设下提供了FedHPD的收敛性理论分析。大量实验表明,FedHPD在多个强化学习基准任务中表现优异,进一步验证了理论结果。此外,额外实验显示,FedHPD无需精心选择公共数据集即可有效运行。

原文摘要 · Abstract (English)

Federated Reinforcement Learning (FedRL) improves sample efficiency while preserving privacy; however, most existing studies assume homogeneous agents, limiting its applicability in real-world scenarios. This paper investigates FedRL in black-box settings with heterogeneous agents, where each agent employs distinct policy networks and training configurations without disclosing their internal details. Knowledge Distillation (KD) is a promising method for facilitating knowledge sharing among heterogeneous models, but it faces challenges related to the scarcity of public datasets and limitations in knowledge representation when applied to FedRL. To address these challenges, we propose Federated Heterogeneous Policy Distillation (FedHPD), which solves the problem of heterogeneous FedRL by utilizing action probability distributions as a medium for knowledge sharing. We provide a theoretical analysis of FedHPD's convergence under standard assumptions. Extensive experiments corroborate that FedHPD shows significant improvements across various reinforcement learning benchmark tasks, further validating our theoretical findings. Moreover, additional experiments demonstrate that FedHPD operates effectively without the need for an elaborate selection of public datasets.

联邦学习强化学习知识蒸馏异构智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。