发现并利用连续强化学习中的值保持结构,提升数据效率与鲁棒性。
Operator-Guided Invariance Learning for Continuous Reinforcement Learning

- 通过李群作用和拉回算子建模值保持映射,发现可迁移的结构规律。
- 在小哈密顿-雅可比-贝尔曼误差下,实现近似值保持结构的严格保证。
- 适合追求高效、稳健连续控制的算法研究者和工业应用开发者。
连续时间与状态/动作空间的强化学习通常数据需求高且对干扰敏感,亟需利用值保持结构来稳定与提升学习性能。现有方法多局限于特定对称性或精确等变性,未能解决如何发现需非线性算子映射的更通用结构。本文提出VPSD-RL(值保持结构发现强化学习),将连续强化学习建模为具有值保持映射的受控扩散过程,该映射由李群作用及关联拉回算子定义。我们证明:当拉回值函数与前推动作运算与受控生成器及奖励泛函可交换时,值保持结构恰好存在;当哈密顿-雅可比-贝尔曼误差较小时,可获得带严格保证的近似结构。该框架通过搜索关联李群算子来发现精确与近似结构。VPSD-RL采用可微分漂移、扩散与奖励模型;通过确定方程残差最小化学习无穷小生成器;利用常微分方程流将其指数化以获得有限变换;并通过状态转移增强与变换一致性正则化集成至连续强化学习中。我们证明:生成器/奖励误差有界时,最优值函数沿近似轨道具有定量稳定性,其敏感度由有效时域决定,并在连续控制基准测试中观察到更高的数据效率与鲁棒性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) with continuous time and state/action spaces is often data-intensive and brittle under nuisance variability and shift, motivating methods that exploit value-preserving structures to stabilize and improve learning. Most existing approaches focus on special cases, such as prescribed symmetries and exact equivariance, without addressing how to discover more general structures that require nonlinear operators to transform and map between continuous state/action systems with isomorphic value functions. We propose \textbf{VPSD-RL} (Value-Preserving Structure Discovery for Reinforcement Learning). It models continuous RL as a controlled diffusion with value-preserving mappings defined through Lie-group actions and associated pullback operators. We show that a value-preserving structure exists exactly when pulling back the value function and pushing forward actions commute with the controlled generator and reward functional. Further, approximate value-preserving structures with rigorous guarantees can be found when the Hamilton--Jacobi--Bellman mismatch is small. This framework discovers exact and approximate value-preserving structures by searching for the associated Lie group operators. VPSD-RL fits differentiable drift, diffusion, and reward models; learns infinitesimal generators via determining-equation residual minimization; exponentiates them with ODE flows to obtain finite transformations; and integrates them into continuous RL through transition augmentation and transformation-consistency regularization. We show that bounded generator/reward mismatch implies quantitative stability of the optimal value function along approximate orbits, with sensitivity governed by the effective horizon, and observe improved data efficiency and robustness on continuous-control benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。