arXiv:2409.03897cs.LGcs.DC2024-09被引 3

联邦Q学习在异构环境中的收敛速度受通信频率影响显著,频繁通信反而拖慢进度。

On the Convergence Rates of Federated Q-Learning across Heterogeneous Environments

  • 通过同步联邦机制让多智能体平均本地Q值估计,每E轮通信一次。
  • 当E>1时,误差收敛速度受限于Θ(E/T),远慢于同质环境。
  • 实验发现收敛存在两阶段现象,分段调学习率可加速整体收敛。

大规模多智能体系统常部署于广泛地理区域,各智能体与异构环境交互。近年来,研究者关注联邦版本的经典强化学习算法中异构性的影响。本文研究同步联邦Q学习,即每E次迭代由K个智能体平均其局部Q估计以学习最优Q函数。我们观察到一个有趣现象:与同质环境类似,采样随机性引起的误差随智能体数K呈线性下降;但与同质情形相反,当E>1时性能显著退化。我们对异构环境下误差演化的精细刻画表明,误差随迭代次数T增加趋近于零。然而,当E>1时收敛缓慢是本质特征而非分析瑕疵。我们证明,对广泛步长范围,误差的ℓ∞范数无法快于Θ(E/T)。此外,实验显示收敛呈现有趣的两阶段现象:给定步长下,初始误差快速下降后反弹并稳定。若能估计相变时间,分阶段使用不同步长可实现更快整体收敛。

原文摘要 · Abstract (English)

Large-scale multi-agent systems are often deployed across wide geographic areas, where agents interact with heterogeneous environments. There is an emerging interest in understanding the role of heterogeneity in the performance of the federated versions of classic reinforcement learning algorithms. In this paper, we study synchronous federated Q-learning, which aims to learn an optimal Q-function by having $K$ agents average their local Q-estimates per $E$ iterations. We observe an interesting phenomenon on the convergence speeds in terms of $K$ and $E$. Similar to the homogeneous environment settings, there is a linear speed-up concerning $K$ in reducing the errors that arise from sampling randomness. Yet, in sharp contrast to the homogeneous settings, $E>1$ leads to significant performance degradation. Specifically, we provide a fine-grained characterization of the error evolution in the presence of environmental heterogeneity, which decay to zero as the number of iterations $T$ increases. The slow convergence of having $E>1$ turns out to be fundamental rather than an artifact of our analysis. We prove that, for a wide range of stepsizes, the $\ell_{\infty}$ norm of the error cannot decay faster than $Θ(E/T)$. In addition, our experiments demonstrate that the convergence exhibits an interesting two-phase phenomenon. For any given stepsize, there is a sharp phase-transition of the convergence: the error decays rapidly in the beginning yet later bounces up and stabilizes. Provided that the phase-transition time can be estimated, choosing different stepsizes for the two phases leads to faster overall convergence.

联邦学习强化学习异构环境收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。