arXiv:2506.15544cs.LG2025-06NeurIPS被引 25

通过稳定梯度流,让深度强化学习在大规模下依然表现稳健。

Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning

  • 分析发现梯度病态与非平稳性是规模化失败主因。
  • 简单干预即可提升不同深度宽度网络的性能稳定性。
  • 适合作为强化学习工程优化的实用方案。

深度强化学习的规模化面临性能退化难题,但其根本原因尚不明确。本文通过一系列实证分析发现,非平稳性与由架构设计不当引发的梯度路径问题共同导致了这一挑战。为此,我们提出一系列直接干预措施,有效稳定梯度传播,使模型在多种网络深度与宽度下均保持鲁棒性能。这些方法实现简单,兼容主流算法,在多个智能体与环境组合上验证了其有效性。

原文摘要 · Abstract (English)

Scaling deep reinforcement learning networks is challenging and often results in degraded performance, yet the root causes of this failure mode remain poorly understood. Several recent works have proposed mechanisms to address this, but they are often complex and fail to highlight the causes underlying this difficulty. In this work, we conduct a series of empirical analyses which suggest that the combination of non-stationarity with gradient pathologies, due to suboptimal architectural choices, underlie the challenges of scale. We propose a series of direct interventions that stabilize gradient flow, enabling robust performance across a range of network depths and widths. Our interventions are simple to implement and compatible with well-established algorithms, and result in an effective mechanism that enables strong performance even at large scales. We validate our findings on a variety of agents and suites of environments.

强化学习梯度稳定大规模训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。