用信任域约束路径空间传输,提升随机最优控制与推断的稳定性
Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference
- 基于信任域的路径空间度量传输,逐步逼近目标分布
- 在扩散采样等任务中显著优于传统方法,提升收敛性能
- 适合需要稳定优化路径的生成模型与控制问题
具有二次控制成本的随机最优控制问题可视为对目标路径空间度量的近似,通常通过梯度优化实现。然而,当目标度量与先验差异较大时,优化过程极具挑战。本文提出一种迭代求解受限问题的方法,引入信任域以系统性地逐步逼近目标度量。该策略可被理解为从先验到目标度量的几何退火过程,其中信任域提供了有原则的退火路径时间步选择方式。我们在多个最优控制应用中验证了该方法的有效性,包括基于扩散的采样、过渡路径采样以及扩散模型微调,结果表明性能显著提升。
原文摘要 · Abstract (English)
Solving stochastic optimal control problems with quadratic control costs can be viewed as approximating a target path space measure, e.g. via gradient-based optimization. In practice, however, this optimization is challenging in particular if the target measure differs substantially from the prior. In this work, we therefore approach the problem by iteratively solving constrained problems incorporating trust regions that aim for approaching the target measure gradually in a systematic way. It turns out that this trust region based strategy can be understood as a geometric annealing from the prior to the target measure, where, however, the incorporated trust regions lead to a principled and educated way of choosing the time steps in the annealing path. We demonstrate in multiple optimal control applications that our novel method can improve performance significantly, including tasks in diffusion-based sampling, transition path sampling, and fine-tuning of diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。