arXiv:2603.17820cs.LG2026-03被引 1

联邦强化学习引入分布式批评者,提升安全关键场景下的决策可靠性。

Federated Distributional Reinforcement Learning with Distributional Critic Regularization

  • 客户端参数化分位值函数,仅联邦化批评者网络
  • 通过局部瓦瑟斯坦中位数约束平均过程,保留分布信息
  • 在多智能体和连续交通环境上显著降低事故率与策略漂移

联邦强化学习通常通过参数平均聚合价值函数或策略,侧重期望回报,可能掩盖安全关键场景中重要的统计多模态和尾部行为。本文提出联邦分布式强化学习(FedDistRL),客户端参数化分位值函数批评者,并仅联邦化这些网络。进一步提出TR-FedDistRL,为每个客户端构建基于时间缓冲区的、风险感知的瓦瑟斯坦中位数。该局部中位数作为参考区域,约束参数平均后的批评者,确保分布信息不被平均稀释。分布信任区域以围绕此参考的收缩-挤压步骤实现。在固定策略评估下,可行性映射为非扩张性,更新在探测集瓦瑟斯坦度量下为压缩性。在老虎机、多智能体网格世界和连续高速公路环境中,相比均值导向及非联邦基线,本方法显著减少均值弥散,改善安全代理指标(灾难/事故率),并降低批评者/策略漂移。

原文摘要 · Abstract (English)

Federated reinforcement learning typically aggregates value functions or policies by parameter averaging, which emphasizes expected return and can obscure statistical multimodality and tail behavior that matter in safety-critical settings. We formalize federated distributional reinforcement learning (FedDistRL), where clients parametrize quantile value function critics and federate these networks only. We also propose TR-FedDistRL, which builds a per client, risk-aware Wasserstein barycenter over a temporal buffer. This local barycenter provides a reference region to constrain the parameter averaged critic, ensuring necessary distributional information is not averaged out during the federation process. The distributional trust region is implemented as a shrink-squash step around this reference. Under fixed-policy evaluation, the feasibility map is nonexpansive and the update is contractive in a probe-set Wasserstein metric under evaluation. Experiments on a bandit, multi-agent gridworld, and continuous highway environment show reduced mean-smearing, improved safety proxies (catastrophe/accident rate), and lower critic/policy drift versus mean-oriented and non-federated baselines.

联邦学习分布强化学习安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。