揭示强化学习中梯度估计噪声比的非均匀性及其对训练不稳定的影响
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
- 通过精确计算噪声信号比,分析REINFORCE算法在特定系统中的梯度估计特性
- 发现随着策略逼近最优解,噪声信号比普遍上升甚至爆炸,导致训练失稳
- 适用于研究策略优化稳定性或设计更鲁棒的强化学习算法的研究者
策略梯度方法广泛应用于强化学习,但训练过程常因学习进展而变得不稳定或变慢。本文通过分析策略梯度估计器的噪声-信号比(NSR),即估计方差(噪声)除以真实梯度范数平方(信号),揭示了这一现象。对于有限时域线性系统与高斯策略、线性状态反馈,以及有限时域多项式系统与高斯策略、多项式反馈,我们可精确刻画REINFORCE估计器的NSR——或以闭式表达,或通过数值矩评估算法实现,无需近似。对于一般非线性动力学和高表达力策略(包括神经网络策略),我们进一步推导出方差的通用上界。这些刻画使我们能直接考察NSR随策略参数的变化及优化轨迹(如SGD、Adam)上的演化规律。在多个实例中,我们发现NSR分布高度非均匀,且通常随策略趋近最优而增加;某些情况下会急剧放大,引发训练不稳和策略坍缩。
原文摘要 · Abstract (English)
Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the noise-to-signal ratio (NSR) of a policy-gradient estimator, defined as the estimator variance (noise) normalized by the squared norm of the true gradient (signal). Our main result is that, for (i) finite-horizon linear systems with Gaussian policies and linear state-feedback, and (ii) finite-horizon polynomial systems with Gaussian policies and polynomial feedback, the NSR of the REINFORCE estimator can be characterized exactly-either in closed form or via numerical moment-evaluation algorithms-without approximation. For general nonlinear dynamics and expressive policies (including neural policies), we further derive a general upper bound on the variance. These characterizations enable a direct examination of how NSR varies across policy parameters and how it evolves along optimization trajectories (e.g. SGD and Adam). Across a range of examples, we find that the NSR landscape is highly non-uniform and typically increases as the policy approaches an optimum; in some regimes it blows up, which can trigger training instability and policy collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。