用优势函数打破虚假关联,提升强化学习泛化能力
Breaking Habits: On the Role of the Advantage Function in Learning Causal State Representations
- 通过优势函数调整动作价值,削弱当前策略下的状态-动作关联
- 在非典型轨迹上性能提升,验证了更优的泛化能力
- 适合关注强化学习因果推理与泛化性的研究者
近期研究表明,强化学习代理可能利用奖励与观测之间的虚假相关性形成策略。这种现象称为策略混淆,源于策略同时影响过去和未来观测变量,形成反馈回路,阻碍代理在常规轨迹外的泛化。本文表明,策略梯度方法中常用的优势函数不仅能降低梯度估计方差,还能缓解策略混淆。通过将动作价值相对于状态表示进行调整,优势函数降低了当前策略下更可能的状态-动作对权重,从而打破虚假关联,促使代理关注因果因素。我们提供了分析与实证证据,证明使用优势函数训练能显著提升跨轨迹表现。
原文摘要 · Abstract (English)
Recent work has shown that reinforcement learning agents can develop policies that exploit spurious correlations between rewards and observations. This phenomenon, known as policy confounding, arises because the agent's policy influences both past and future observation variables, creating a feedback loop that can hinder the agent's ability to generalize beyond its usual trajectories. In this paper, we show that the advantage function, commonly used in policy gradient methods, not only reduces the variance of gradient estimates but also mitigates the effects of policy confounding. By adjusting action values relative to the state representation, the advantage function downweights state-action pairs that are more likely under the current policy, breaking spurious correlations and encouraging the agent to focus on causal factors. We provide both analytical and empirical evidence demonstrating that training with the advantage function leads to improved out-of-trajectory performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。