提出对转移度量进行预移位,使强化学习中的低秩结构更易捕捉。
Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning
- 通过预移位转移度量,使其自然呈现低秩特性
- 理论证明移位后可实现高精度低秩估计,误差由谱可恢复性决定
- 实验验证移位策略能提升目标条件强化学习性能
低秩结构是许多现代强化学习算法的隐含假设。例如,无奖励和目标条件强化学习方法通常假设后续度量具有低秩表示。本文指出,原始后续度量并非近似低秩;相反,经过若干初始转移跳过后,其移位版本会自然呈现低秩结构。我们为从采样条目中估计该移位度量的低秩近似提供了有限样本性能保证。分析表明,近似与估计误差主要受新提出的矩阵谱可恢复性参数控制。为此,我们推导出一类新的马尔可夫链函数不等式——Ⅱ型泊松不等式,从而量化有效低秩近似所需的移位量。分析显示,所需移位量取决于移位后后续度量高阶奇异值的衰减速率,在实践中通常较小。此外,我们建立了移位量与系统局部混合性质之间的联系,为移位选择提供自然依据。最后,实验验证了理论发现,证明移位后的后续度量确实能提升目标条件强化学习性能。
原文摘要 · Abstract (English)
Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume that the successor measure admits a low-rank representation. In this work, we challenge this assumption by first remarking that the successor measure itself is not approximately low-rank. Instead, we demonstrate that a low-rank structure naturally emerges in the shifted successor measure, which captures the system dynamics after bypassing a few initial transitions. We provide finite-sample performance guarantees for the entry-wise estimation of a low-rank approximation of the shifted successor measure from sampled entries. Our analysis reveals that both the approximation and estimation errors are primarily governed by a newly introduced quantitity: the spectral recoverability of the corresponding matrix. To bound this parameter, we derive a new class of functional inequalities for Markov chains that we call Type II Poincaré inequalities and from which we can quantify the amount of shift needed for effective low-rank approximation and estimation. This analysis shows in particular that the required shift depends on decay of the high-order singular values of the shifted successor measure and is hence typically small in practice. Additionally, we establish a connection between the necessary shift and the local mixing properties of the underlying dynamical system, which provides a natural way of selecting the shift. Finally, we validate our theoretical findings with experiments, and demonstrate that shifting the successor measure indeed leads to improved performance in goal-conditioned RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。