解释强化学习中为何用奖励累积代替全程回报
On the "Causality" Step in Policy Gradient Derivations: A Pedagogical Reconciliation of Full Return and Reward-to-Go
- 通过前缀轨迹分布与得分函数恒等式,严谨推导奖励累积的来源
- 证明奖励累积与全程回报在数学上等价,不影响估计器性能
- 适合想搞懂强化学习理论细节的研究者阅读
在强化学习政策梯度的入门讲解中,通常先用完整轨迹回报推导REINFORCE估计器,再以‘因果性’为由替换为奖励累积。尽管结论正确,但该步骤常缺乏严格推导,导致过去奖励项去向不明。本文明确分离这一关键步骤,基于前缀轨迹分布与得分函数恒等式,给出数学上严格的推导。结果表明,奖励累积并非对全程回报的事后修正,而是目标函数按前缀轨迹分解后的自然产物。在此框架下,传统因果性论证成为推导的推论,而非额外的启发式原则。
原文摘要 · Abstract (English)
In introductory presentations of policy gradients, one often derives the REINFORCE estimator using the full trajectory return and then states, by ``causality,'' that the full return may be replaced by the reward-to-go. Although this statement is correct, it is frequently presented at a level of rigor that leaves unclear where the past-reward terms disappear. This short paper isolates that step and gives a mathematically explicit derivation based on prefix trajectory distributions and the score-function identity. The resulting account does not change the estimator. Its contribution is conceptual: instead of presenting reward-to-go as a post hoc unbiased replacement for full return, it shows that reward-to-go arises directly once the objective is decomposed over prefix trajectories. In this formulation, the usual causality argument is recovered as a corollary of the derivation rather than as an additional heuristic principle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。