arXiv:2505.23459cs.LG2025-05被引 2

提出联邦强化学习的收敛速率,揭示异质环境下策略需随机化。

On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments

  • 用投影类算子分析联邦平均的非凸优化,突破传统假设
  • 首次给出带熵正则的联邦策略梯度显式收敛常数
  • 证明异质环境迫使策略必须随机,与单智能体不同

我们给出了无熵正则和带熵正则的联邦软最大化随机策略梯度(FedPG)在本地训练下的全局收敛速率。结果表明,FedPG 能以平均智能体价值接近最优,其差距由异质性水平控制。值得注意的是,我们首次获得带熵正则策略梯度的显式收敛速率,并引入投影类算子实现。分析基于对联邦平均在非凸目标下的新理解:单智能体设定中的 Łojasiewicz 型不等式(Mei 等,2020)在联邦目标下不成立。这揭示了单智能体与联邦强化学习的根本差异——单智能体最优策略可为确定性,而联邦目标可能固有地要求随机策略。

原文摘要 · Abstract (English)

We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient with explicit constants, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the Łojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies.

联邦学习强化学习策略梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。