提出无需调参的联邦TD学习方法,可在异构环境中实现最优收敛速度。
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
- 采用双时间尺度与Polyak-Ruppert平均,实现参数无关的联邦TD学习。
- 在平均奖励和折扣设定下,均达到最优收敛率$ ilde{O}(1/NT)$。
- 适用于异构环境下的联邦强化学习,适合实际多智能体系统部署。
联邦学习(FL)可通过在多个智能体间分布探索与训练,显著加速强化学习。现有方法在独立同分布数据下可实现$ ilde{O}(1/(NT))$的最优收敛率,其中$T$为迭代次数,$N$为智能体数。然而,当数据来自马尔可夫链时,现有TD学习算法需依赖未知问题参数才能达到该速率。本文提出一种双时间尺度联邦时序差分(FTD)学习方法,结合Polyak-Ruppert平均,首次在平均奖励和折扣设定下,实现无需调参的$ ilde{O}(1/NT)$最优收敛率,适用于异构环境中的联邦学习。尽管结果在单智能体场景亦具新意,但其更适用于真实且更具挑战性的异构联邦强化学习场景。
原文摘要 · Abstract (English)
Federated learning (FL) can dramatically speed up reinforcement learning by distributing exploration and training across multiple agents. It can guarantee an optimal convergence rate that scales linearly in the number of agents, i.e., a rate of $\tilde{O}(1/(NT)),$ where $T$ is the iteration index and $N$ is the number of agents. However, when the training samples arise from a Markov chain, existing results on TD learning achieving this rate require the algorithm to depend on unknown problem parameters. We close this gap by proposing a two-timescale Federated Temporal Difference (FTD) learning with Polyak-Ruppert averaging. Our method provably attains the optimal $\tilde{O}(1/NT)$ rate in both average-reward and discounted settings--offering a parameter-free FTD approach for Markovian data. Although our results are novel even in the single-agent setting, they apply to the more realistic and challenging scenario of FL with heterogeneous environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。