通过贝尔曼算子与特征协方差的谱关系,统一了强化学习中的表征学习与探索。
Spectral Bellman Method: Unifying Representation and Exploration in RL
- 基于贝尔曼误差为零时的谱特性,建立价值函数分布变换与特征协方差的内在联系。
- 在困难探索和长周期任务中,新表征显著提升智能体性能。
- 适用于需结构化探索的价值型强化学习,可无缝集成至现有算法。
表征学习对强化学习的实证与理论成功至关重要。然而,许多现有方法源于模型学习视角,与实际强化学习任务不匹配。本文提出谱贝尔曼方法(Spectral Bellman Method),从固有贝尔曼误差(IBE)条件出发,将表征学习与贝尔曼更新在价值函数空间中的基本结构对齐,直接适配基于价值的强化学习。核心洞见在于:在零IBE条件下,贝尔曼算子对价值函数分布的变换与特征协方差结构存在本质谱关联。这一联系催生了一种新的、理论驱动的表征学习目标,仅需对现有算法进行简单修改即可实现。我们证明,所学表征能通过将特征协方差与贝尔曼动力学对齐,实现结构化探索,在硬探索和长周期任务中表现更优。该框架自然扩展至多步贝尔曼算子,为构建更强大且结构严谨的价值型强化学习表征提供了原则性路径。
原文摘要 · Abstract (English)
Representation learning is critical to the empirical and theoretical success of reinforcement learning. However, many existing methods are induced from model-learning aspects, misaligning them with the RL task in hand. This work introduces the Spectral Bellman Method, a novel framework derived from the Inherent Bellman Error (IBE) condition. It aligns representation learning with the fundamental structure of Bellman updates across a \textit{space} of possible value functions, making it directly suited for value-based RL. Our key insight is a fundamental spectral relationship: under the zero-IBE condition, the transformation of a \textit{distribution} of value functions by the Bellman operator is intrinsically linked to the feature covariance structure. This connection yields a new, theoretically-grounded objective for learning state-action features that capture this Bellman-aligned covariance, requiring only a simple modification to existing algorithms. We demonstrate that our learned representations enable structured exploration by aligning feature covariance with Bellman dynamics, improving performance in hard-exploration and long-horizon tasks. Our framework naturally extends to multi-step Bellman operators, offering a principled path toward learning more powerful and structurally sound representations for value-based RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。