揭示策略梯度方法中分布偏差的影响及其缓解机制
Analysis of On-policy Policy Gradient Methods under the Distribution Mismatch
- 分析了策略梯度中因分布不匹配导致的梯度偏差问题
- 证明在折扣因子趋近1时,偏差与分布误差均线性收敛至0
- 为实际算法与理论框架之间的差距提供解释,适合强化学习研究者
策略梯度方法是解决复杂强化学习问题最成功的方法之一。尽管在实践中表现优异,但许多针对折扣问题的先进策略梯度算法因存在分布不匹配而偏离了理论上的策略梯度定理。本文分析了这种不匹配对策略梯度方法的影响。首先,在表格参数化情形下,我们证明由不匹配引起的有偏梯度仍能有效表征全局最优性。随后,通过在广义参数化下推导状态分布不匹配和梯度不匹配的显式界,发现这些界在折扣因子趋近于1时至少以线性速度收敛至0。基于这些界,我们进一步建立了有偏策略梯度迭代的保证,表明其会趋近于相对于精确梯度的近似驻点,渐近残差依赖于折扣因子。研究成果揭示了策略梯度方法的鲁棒性,并解释了理论基础与实际实现间的差距。
原文摘要 · Abstract (English)
Policy gradient methods are one of the most successful approaches for solving challenging reinforcement learning problems. Despite their empirical successes, many state-of-the-art policy gradient algorithms for discounted problems deviate from the theoretical policy gradient theorem due to the existence of a distribution mismatch. In this work, we analyze the impact of this mismatch on policy gradient methods. Specifically, we first show that in the case of tabular parameterizations, the biased gradient induced by the mismatch still yields a valid first-order characterization of global optimality. Then, we extend this analysis to more general parameterizations by deriving explicit bounds on both the state distribution mismatch and the resulting gradient mismatch in episodic and continuing MDPs, which are shown to vanish at least linearly as the discount factor approaches one. Building on these bounds, we further establish guarantees for the biased policy gradient iterates, showing that they approach approximate stationary points with respect to the exact gradient, with asymptotic residuals depending on the discount factor. Our findings offer insights into the robustness of policy gradient methods as well as the gap between theoretical foundations and practical implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。