arXiv:2509.24305cs.LGcs.DC2025-09

提出异步策略梯度聚合算法,提升分布式强化学习效率。

Asynchronous Policy Gradient Aggregation for Efficient Distributed Reinforcement Learning

  • 设计异步梯度聚合机制,支持非同步计算与通信。
  • 在同质和异质环境下均实现更优理论复杂度与性能。
  • 适合大规模分布式强化学习系统,尤其在异构环境中表现优异。

我们研究了在异步并行计算与通信条件下,基于策略梯度方法的分布式强化学习。尽管非分布式方法在理论上已较为成熟且取得显著实证成果,其分布式版本仍缺乏深入探索,尤其是在存在异构异步计算和通信瓶颈的情况下。本文提出两种新算法:Rennala NIGT 与 Malenia NIGT,实现异步策略梯度聚合,并达到当前最优效率。在同质设置下,Rennala NIGT 可证明地降低总计算与通信复杂度,同时支持 AllReduce 操作;在异质设置下,Malenia NIGT 同时处理异步计算与异构环境,具有更优的理论保证。实验结果进一步验证了其显著优于现有方法的性能。

原文摘要 · Abstract (English)

We study distributed reinforcement learning (RL) with policy gradient methods under asynchronous and parallel computations and communications. While non-distributed methods are well understood theoretically and have achieved remarkable empirical success, their distributed counterparts remain less explored, particularly in the presence of heterogeneous asynchronous computations and communication bottlenecks. We introduce two new algorithms, Rennala NIGT and Malenia NIGT, which implement asynchronous policy gradient aggregation and achieve state-of-the-art efficiency. In the homogeneous setting, Rennala NIGT provably improves the total computational and communication complexity while supporting the AllReduce operation. In the heterogeneous setting, Malenia NIGT simultaneously handles asynchronous computations and heterogeneous environments with strictly better theoretical guarantees. Our results are further corroborated by experiments, showing that our methods significantly outperform prior approaches.

强化学习分布式异步梯度聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。