arXiv:2603.23173cs.LGmath.OC2026-03中稿 · ICLR

用量子力学方法解决长时域随机控制难题,精度提升10倍且效率更高。

A Schrödinger Eigenfunction Method for Long-Horizon Stochastic Optimal Control

  • 将控制问题转化为量子薛定谔算子的本征系统求解
  • 在对称LQR中获得任意终值代价的解析解,长时域控制精度大幅提升
  • 提出新损失函数克服神经网络学习中的隐式重加权问题,适合高维长时控制

高维随机最优控制(SOC)随规划时长增加而急剧变难:现有方法在时长 $T$ 上线性增长,性能常指数下降。本文针对一类可解的线性可解SOC问题——其无控漂移为势能梯度的情形——实现突破。在此设定下,哈密顿-雅可比-贝尔曼方程简化为由算子 $/mathcal{L}$ 控制的线性偏微分方程。我们证明,在梯度漂移假设下,$/mathcal{L}$ 与薛定谔算子 $/mathcal{S} = -Δ+ /mathcal{V}$ 单位等价,具有离散谱,使长时域控制可通过 $/mathcal{L}$ 的本征系统高效描述。由此得出两项关键成果:第一,对称线性二次调节器(LQR)中,$/mathcal{S}$ 对应量子谐振子哈密顿量,其闭式本征系统给出任意终值代价下的解析解;第二,在更一般情形下,使用神经网络学习 $/mathcal{L}$ 的本征系统。我们识别出现有本征函数学习损失的隐式重加权问题会降低控制性能,并提出新损失函数加以缓解。在多个长时域基准测试中,该方法相较最先进方法控制精度提升一个数量级,同时内存与运行复杂度从 $/mathcal{O}(Td)$ 降至 $/mathcal{O}(d)$。

原文摘要 · Abstract (English)

High-dimensional stochastic optimal control (SOC) becomes harder with longer planning horizons: existing methods scale linearly in the horizon $T$, with performance often deteriorating exponentially. We overcome these limitations for a subclass of linearly-solvable SOC problems-those whose uncontrolled drift is the gradient of a potential. In this setting, the Hamilton-Jacobi-Bellman equation reduces to a linear PDE governed by an operator $\mathcal{L}$. We prove that, under the gradient drift assumption, $\mathcal{L}$ is unitarily equivalent to a Schrödinger operator $\mathcal{S} = -Δ+ \mathcal{V}$ with purely discrete spectrum, allowing the long-horizon control to be efficiently described via the eigensystem of $\mathcal{L}$. This connection provides two key results: first, for a symmetric linear-quadratic regulator (LQR), $\mathcal{S}$ matches the Hamiltonian of a quantum harmonic oscillator, whose closed-form eigensystem yields an analytic solution to the symmetric LQR with \emph{arbitrary} terminal cost. Second, in a more general setting, we learn the eigensystem of $\mathcal{L}$ using neural networks. We identify implicit reweighting issues with existing eigenfunction learning losses that degrade performance in control tasks, and propose a novel loss function to mitigate this. We evaluate our method on several long-horizon benchmarks, achieving an order-of-magnitude improvement in control accuracy compared to state-of-the-art methods, while reducing memory usage and runtime complexity from $\mathcal{O}(Td)$ to $\mathcal{O}(d)$.

随机控制量子类比神经网络长时域优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。