arXiv:2410.05163cs.LGmath.OC2024-10被引 4

提出高效求解随机最优控制的新方法,显著提升计算速度与内存效率。

An Efficient On-Policy Deep Learning Framework for Stochastic Optimal Control

  • 基于吉尔萨诺夫定理直接计算策略梯度,避免复杂反向传播
  • 在高维与长时序问题上表现优异,计算速度与内存消耗大幅降低
  • 适用于分布采样和扩散模型微调,适合需要高效训练的控制任务

我们提出一种新型的在线策略算法,用于求解随机最优控制(SOC)问题。通过利用吉尔萨诺夫定理,该方法无需对随机微分方程进行昂贵的反向传播或求解伴随问题,即可直接计算在线策略梯度。这一方法显著加速了神经网络控制策略的优化过程,并在高维问题与长时域场景下具有良好的可扩展性。我们在经典SOC基准上进行了评估,同时应用于通过施罗丁格-福尔默过程从非归一化分布采样以及微调预训练扩散模型。实验结果表明,相较于现有方法,该方法在计算速度和内存效率方面均有显著提升。

原文摘要 · Abstract (English)

We present a novel on-policy algorithm for solving stochastic optimal control (SOC) problems. By leveraging the Girsanov theorem, our method directly computes on-policy gradients of the SOC objective without expensive backpropagation through stochastic differential equations or adjoint problem solutions. This approach significantly accelerates the optimization of neural network control policies while scaling efficiently to high-dimensional problems and long time horizons. We evaluate our method on classical SOC benchmarks as well as applications to sampling from unnormalized distributions via Schrödinger-Föllmer processes and fine-tuning pre-trained diffusion models. Experimental results demonstrate substantial improvements in both computational speed and memory efficiency compared to existing approaches.

随机控制深度学习高效算法扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。