揭示状态图连通性如何影响强化学习中的谱表示精度。
Impact of Connectivity on Laplacian Representations in Reinforcement Learning
- 用图拉普拉斯特征向量构建状态表示,结合采样轨迹估计谱特征。
- 证明近似误差随状态图代数连通性增大而减小,理论可量化。
- 适用于任意非均匀策略,对复杂环境中的表示学习有指导意义。
在大规模强化学习中,学习紧凑的状态表示对于缓解维度灾难至关重要。现有方法通过构建状态表示为状态图拉普拉斯特征向量的线性组合,利用马尔可夫决策过程(MDP)的结构先验。当转移图未知或状态空间过大时,可通过样本轨迹直接估计图谱特征。本文证明了基于学习到的谱特征进行线性值函数近似的上界误差,并表明该误差随状态图的代数连通性变化而变化,将近似质量与MDP的拓扑结构联系起来。进一步,我们给出了特征向量估计引入误差的边界,实现了表示学习全流程的端到端误差分解。此外,本文给出的强化学习场景下拉普拉斯算子表达式虽与已有等价,但能避免文献中一些常见误解。理论结果适用于一般(非均匀)策略,无需对诱导转移核的对称性做假设。我们在网格世界环境中通过数值模拟验证了理论发现。
原文摘要 · Abstract (English)
Learning compact state representations in Markov Decision Processes (MDPs) has proven crucial for addressing the curse of dimensionality in large-scale reinforcement learning (RL) problems. Existing principled approaches leverage structural priors on the MDP by constructing state representations as linear combinations of the state-graph Laplacian eigenvectors. When the transition graph is unknown or the state space is prohibitively large, the graph spectral features can be estimated directly via sample trajectories. In this work, we prove an upper bound on the approximation error of linear value function approximation under the learned spectral features. We show how this error scales with the algebraic connectivity of the state-graph, grounding the approximation quality in the topological structure of the MDP. We further bound the error introduced by the eigenvector estimation itself, leading to an end-to-end error decomposition across the representation learning pipeline. Additionally, our expression of the Laplacian operator for the RL setting, although equivalent to existing ones, prevents some common misunderstandings, of which we show some examples from the literature. Our results hold for general (non-uniform) policies without any assumptions on the symmetry of the induced transition kernel. We validate our theoretical findings with numerical simulations on gridworld environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。