arXiv:2609.04005cs.AI2026-09

用双重平坦几何重构强化学习规划,让决策更高效。

The Dually Flat Geometry of Planning as Inference

  • 将规划嵌入动态过程,定义新状态访问度量
  • 访问度量构成双重平坦流形,支持非线性奖励优化
  • 适合研究强化学习与神经科学的理论工作者

我们提出一种新的强化学习占用度量表征,通过重置规划过程将规划准则嵌入动态系统。其稳态分布称为访问度量,是决策信息几何最自然的表达对象。可实现的访问度量构成一个双重平坦统计流形,其两个仿射坐标分别为访问概率和对数策略,二者在条件熵下对偶。该结构使规划-推断方法能从线性奖励推广到访问度量的非线性泛函,每次迭代仅需一步自然梯度更新,并赋予时序差分误差以边际效用估计的解释。本文发展了该几何及其在强化学习与理论神经科学中的影响。

原文摘要 · Abstract (English)

We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.

强化学习信息几何规划推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。