arXiv:2507.14487cs.LG2025-07被引 1

解决异构环境下的联邦强化学习,提升全局策略的鲁棒性。

Federated Reinforcement Learning in Heterogeneous Environments

  • 设计新目标函数,让全局策略在不同环境中表现稳定。
  • 提出FedRQ算法,理论证明可收敛到最优策略。
  • 适配连续状态空间,兼容主流深度强化学习模型。

我们研究了存在环境异构性的联邦强化学习(FRL-EH)框架,其中各本地环境呈现统计异构性。在此框架下,智能体通过聚合集体经验协作学习全局策略,同时保护本地轨迹隐私。为更贴近真实场景,提出一种鲁棒的FRL-EH框架,引入新型全局目标函数,旨在优化一个在异构本地环境及其合理扰动下均表现稳健的全局策略。我们提出一种基于表格的FRL算法FedRQ,并理论证明其渐近收敛至全局目标函数的最优策略。此外,通过使用期望损失(expectile loss),将FedRQ扩展至连续状态空间,解决了在连续状态子集上最小化价值函数的关键挑战。这一进展使FedRQ原则可无缝融入多种基于深度神经网络(DNN)的强化学习算法。大量实验评估验证了所提FRL算法在多样化异构环境中的有效性和鲁棒性,性能持续优于现有最先进方法。

原文摘要 · Abstract (English)

We investigate a Federated Reinforcement Learning with Environment Heterogeneity (FRL-EH) framework, where local environments exhibit statistical heterogeneity. Within this framework, agents collaboratively learn a global policy by aggregating their collective experiences while preserving the privacy of their local trajectories. To better reflect real-world scenarios, we introduce a robust FRL-EH framework by presenting a novel global objective function. This function is specifically designed to optimize a global policy that ensures robust performance across heterogeneous local environments and their plausible perturbations. We propose a tabular FRL algorithm named FedRQ and theoretically prove its asymptotic convergence to an optimal policy for the global objective function. Furthermore, we extend FedRQ to environments with continuous state space through the use of expectile loss, addressing the key challenge of minimizing a value function over a continuous subset of the state space. This advancement facilitates the seamless integration of the principles of FedRQ with various Deep Neural Network (DNN)-based RL algorithms. Extensive empirical evaluations validate the effectiveness and robustness of our FRL algorithms across diverse heterogeneous environments, consistently achieving superior performance over the existing state-of-the-art FRL algorithms.

联邦学习强化学习异构环境策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。