arXiv:2504.15753math.OCcs.LG2025-04被引 3

提出可解的非高斯分布调控方法,突破传统扩散模型限制。

Markov Kernels, Distances and Optimal Control: A Parable of Linear Quadratic Non-Gaussian Distribution Steering

  • 通过最优控制构造时变距离函数,推导新型马尔可夫核。
  • 首次实现线性二次非高斯薛定谔桥的精确求解。
  • 适合研究随机控制与生成模型交叉领域的学者。

针对可控线性时变系统(A_t, B_t)及半正定矩阵 Q_t,本文推导了Itô扩散过程 dx_t = A_t x_t dt + √2 B_t dw_t 搭配以速率 (1/2) x^T Q_t x 进行概率质量衰减的马尔可夫核。该核为对应线性反应-对流-扩散偏微分方程的格林函数。结果推广了此前 (A_t, B_t) = (0, I) 的特例,并依赖于关联Riccati矩阵常微分方程的解。此结果表明,线性二次非高斯薛定谔桥可精确求解:在固定时间窗口内,将受控线性时变扩散从给定非高斯分布引导至目标分布,同时最小化期望二次代价,可通过动态Sinkhorn递推结合所推导核实现。新方法基于求解确定性最优控制问题获得状态-时间依赖的距离型泛函,突破了以往依赖赫尔米特多项式或威爾衍算等方法的局限。该技术揭示了马尔可夫核、距离与最优控制间的深层联系,意义超越其在薛定谔桥中的应用。

原文摘要 · Abstract (English)

For a controllable linear time-varying (LTV) pair $(\boldsymbol{A}_t,\boldsymbol{B}_t)$ and $\boldsymbol{Q}_{t}$ positive semidefinite, we derive the Markov kernel for the Itô diffusion ${\mathrm{d}}\boldsymbol{x}_{t}=\boldsymbol{A}_{t}\boldsymbol{x}_t {\mathrm{d}} t + \sqrt{2}\boldsymbol{B}_{t}{\mathrm{d}}\boldsymbol{w}_{t}$ with an accompanying killing of probability mass at rate $\frac{1}{2}\boldsymbol{x}^{\top}\boldsymbol{Q}_{t}\boldsymbol{x}$. This Markov kernel is the Green's function for an associated linear reaction-advection-diffusion partial differential equation. Our result generalizes the recently derived kernel for the special case $\left(\boldsymbol{A}_t,\boldsymbol{B}_t\right)=\left(\boldsymbol{0},\boldsymbol{I}\right)$, and depends on the solution of an associated Riccati matrix ODE. A consequence of this result is that the linear quadratic non-Gaussian Schrödinger bridge is exactly solvable. This means that the problem of steering a controlled LTV diffusion from a given non-Gaussian distribution to another over a fixed deadline while minimizing an expected quadratic cost can be solved using dynamic Sinkhorn recursions performed with the derived kernel. Our derivation for the $\left(\boldsymbol{A}_t,\boldsymbol{B}_t,\boldsymbol{Q}_t\right)$-parametrized kernel pursues a new idea that relies on finding a state-time dependent distance-like functional given by the solution of a deterministic optimal control problem. This technique breaks away from existing methods, such as generalizing Hermite polynomials or Weyl calculus, which have seen limited success in the reaction-diffusion context. Our technique uncovers a new connection between Markov kernels, distances, and optimal control. This connection is of interest beyond its immediate application in solving the linear quadratic Schrödinger bridge problem.

最优控制扩散模型非高斯薛定谔桥

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。