arXiv:2512.17688cs.LGstat.ML2025-12

首次证明异构联邦SARSA在本地训练下可收敛并实现线性加速。

Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents

  • 提出多步误差展开新方法,精准分析联邦强化学习误差来源。
  • 在异构环境与多轮本地训练下,证明算法仍能收敛且通信复杂度可控。
  • 适合研究联邦强化学习理论或分布式智能体系统的设计者。

我们为采用线性函数近似和本地训练的联邦SARSA(FedSARSA)提供了新颖的理论分析。在局部转移与奖励存在异构性的条件下,建立了首个该场景下的样本复杂度与通信复杂度边界。分析核心是一个针对单智能体SARSA的新颖精确多步误差展开,具有独立研究价值。我们的分析准确量化了异构性的影响,证明了在多轮本地更新下FedSARSA仍可收敛。关键的是,我们表明在智能体数量上,FedSARSA可实现线性加速,仅受马尔可夫采样带来的高阶项限制。数值实验验证了理论结论。

原文摘要 · Abstract (English)

We present a novel theoretical analysis of Federated SARSA (FedSARSA) with linear function approximation and local training. We establish convergence guarantees for FedSARSA in the presence of heterogeneity, both in local transitions and rewards, providing the first sample and communication complexity bounds in this setting. At the core of our analysis is a new, exact multi-step error expansion for single-agent SARSA, which is of independent interest. Our analysis precisely quantifies the impact of heterogeneity, demonstrating the convergence of FedSARSA with multiple local updates. Crucially, we show that FedSARSA achieves linear speed-up with respect to the number of agents, up to higher-order terms due to Markovian sampling. Numerical experiments support our theoretical findings.

联邦学习强化学习收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。