arXiv:2502.20476cs.LG2025-02被引 3

统一了三种控制方法的优化本质,揭示其共享梯度更新机制。

Unifying Model Predictive Path Integral Control, Reinforcement Learning, and Diffusion Models for Optimal Control and Planning

  • 从吉布斯测度出发,用梯度优化统一三类方法。
  • 证明策略梯度可转为MPPI,扩散模型反向采样与之同构。
  • 适合研究最优控制、强化学习与生成模型交叉的学者。

模型预测路径积分(MPPI)控制、强化学习(RL)和扩散模型在轨迹优化、决策与运动规划中均表现出色,但传统上被视为独立的方法。本文通过在吉布斯测度上的梯度优化,建立三者统一视角。首先,证明MPPI等价于对平滑能量函数进行梯度上升;其次,表明策略梯度方法经指数变换后可化为MPPI;此外,还证实扩散模型的逆向采样过程遵循与MPPI相同的更新规则。

原文摘要 · Abstract (English)

Model Predictive Path Integral (MPPI) control, Reinforcement Learning (RL), and Diffusion Models have each demonstrated strong performance in trajectory optimization, decision-making, and motion planning. However, these approaches have traditionally been treated as distinct methodologies with separate optimization frameworks. In this work, we establish a unified perspective that connects MPPI, RL, and Diffusion Models through gradient-based optimization on the Gibbs measure. We first show that MPPI can be interpreted as performing gradient ascent on a smoothed energy function. We then demonstrate that Policy Gradient methods reduce to MPPI by applying an exponential transformation to the objective function. Additionally, we establish that the reverse sampling process in diffusion models follows the same update rule as MPPI.

最优控制强化学习扩散模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。