arXiv:2605.27385cs.LGcs.AI2026-05中稿 · the International …

针对异构环境下的联邦强化学习,提出个性化观测归一化方法。

Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity

论文配图:Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity
图 1 · 摘自论文原文
  • 每智能体本地用动态均值方差归一化观测输入
  • 相比基线方法训练更快,性能更优
  • 适合存在数据分布差异的联邦强化学习场景

联邦强化学习(FedRL)允许多个智能体在不共享原始数据的前提下协同训练全局策略,适用于隐私敏感场景。然而,在状态转移动态不同的异构环境中,输入分布不一致导致参数更新不平衡。为此,本文提出个性化观测归一化(PON)方法,使各智能体本地使用持续更新的运行均值和方差对原始状态输入进行归一化,确保局部特征缩放一致且聚合时不被掩盖。实验表明,跨智能体共享归一化参数无效,因本地输入分布差异显著,凸显个性化统计的必要性。在异构MuJoCo任务上的测试显示,所提PON方法加速了训练并优于基线方法。

原文摘要 · Abstract (English)

Federated reinforcement learning (FedRL) enables multiple agents to collaboratively train a global policy without sharing raw data, making it ideal for privacy-sensitive applications. However, FedRL faces challenges in heterogeneous environments where differing state-transition dynamics lead to non-identical input distributions and imbalanced parameter updates during aggregation. Therefore, this paper develops a personalized observation normalization (PON) method, allowing each agent to locally normalize raw state inputs using a continuously updated running mean and variance. This design ensures consistent scaling of local feature without overshadowing across agents during aggregation. Furthermore, we demonstrate that sharing normalization parameters across agents is ineffective due to the diverse local input distributions, which highlights the necessity of personalized statistics. Experiments on heterogeneous MuJoCo tasks show that our developed PON accelerates training and achieves superior performance compared to baseline methods.

联邦学习强化学习归一化异构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。