用最优传输理论重新定义强化学习策略优化,为算法设计提供几何新视角。
Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization

- 将策略视为动作概率分布的映射,构建基于平稳分布的黎曼结构。
- 推导出能量函数的梯度与海森矩阵,实现二阶优化分析。
- 适用于低维问题精确求解,高维时结合神经网络进行近似优化。
我们提出一种强化学习(RL)的几何框架,将策略视为映射到动作概率分布的Wasserstein空间。首先,在一般条件下证明了由平稳分布诱导的黎曼结构的存在性。接着,定义了策略的切空间并刻画了测地线,特别关注从状态空间到动作空间概率测度切空间的向量场的可测性。然后,构建了一般RL优化问题,并利用Otto微积分构造梯度流。计算了能量函数的梯度与海森矩阵,实现了形式化的二阶分析。最后,通过低维问题的数值例子验证方法,直接从理论框架中计算梯度;对于高维问题,采用神经网络参数化策略,并基于成本的遍历近似进行优化。
原文摘要 · Abstract (English)
We present a geometric framework for Reinforcement Learning (RL) that views policies as maps into the Wasserstein space of action probabilities. First, we define a Riemannian structure induced by stationary distributions, proving its existence in a general context. We then define the tangent space of policies and characterize the geodesics, specifically addressing the measurability of vector fields mapped from the state space to the tangent space of probability measures over the action space. Next, we formulate a general RL optimization problem and construct a gradient flow using Otto's calculus. We compute the gradient and the Hessian of the energy, providing a formal second-order analysis. Finally, we illustrate the method with numerical examples for low-dimensional problems, computing the gradient directly from our theoretical formalism. For high-dimensional problems, we parameterize the policy using a neural network and optimize it based on an ergodic approximation of the cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。