arXiv:2606.18186cs.LGcs.AI2026-06

用偏微分方程改进扩散策略,提升长时序控制的稳定性和可靠性。

Kolmogorov Regression for Robust Diffusion Policies

  • 将随机得分匹配转为确定性边界值问题,基于高斯测度理论设计新损失函数。
  • 在推断中实现67.6%的步间漂移减少,且零奖励下仍可检测失败。
  • 适用于机器人操控与制造调度,特别适合对稳定性要求高的场景。

有限维扩散策略因离散化误差导致时间漂移,影响长时序任务表现。本文提出一种反向柯尔莫哥洛夫方程,将扩散策略提升至卡梅隆-马丁空间(希尔伯特空间子集),以确定性边值问题替代随机得分匹配。核心创新基于高斯测度理论,通过彩色噪声分布定义模型采样样本的正则性。训练采用导出的精度加权卡梅隆-马丁损失,并引入柯尔莫哥洛夫残差作为推理阶段的PDE诊断工具。该方法实现:(i) 收敛性保证,界中常数依赖于核的有效秩而非动作维度;(ii) 通过谱加权提升轨迹平滑性;(iii) 无需奖励信号即可实现确定性故障检测。在两个应用领域验证:在PushT操作任务中,卡梅隆-马丁损失使最大回合奖励提升17%(0.95 vs. 0.78),推理期间步间漂移减少67.6%;在6站连续生产线上,相比经典LSTM基线,RMSE降低28.4%,测试周期内饥饿事件召回率达1.0,瓶颈识别准确率(Precision@1)达1.0,信噪比提升13倍。进一步通过哈密顿-雅可比可达性理论认证调度策略,在100次模拟中将死锁事件减少96%(防止351次事件)。

原文摘要 · Abstract (English)

Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (when deployed on physical systems). We introduce a backward Kolmogorov equation that lifts diffusion policies to a Cameron-Martin space -- a subset of the Hilbert space. Essentially, replacing stochastic score matching with a deterministic boundary-value PDE problem. Our core innovation thrives on Gaussian measure theory whereupon the diffusion noise covariance operator is realized from a colored noise distribution which prescribes a notion of regularity on samples from the model at inference time. We train the diffusion model with a derived precision-weighted Cameron- Martin loss and a Kolmogorov residual is introduced as a PDE diagnostic during inference. These substitutions yield (i) convergence guarantees where the bound's constants depend on the effective rank of the kernel rather than action dimension, (ii) improved trajectory regularity via spectral weighting, and (iii) a deterministic failure detector without reward signals. Validation across two application domains demonstrates substantial improvements: on the PushT manipulation benchmark, the Cameron-Martin loss achieves a 17% improvement in maximum episode reward (0.95 vs. 0.78 for MSE) and 67.6% reduction in inter-step drifts during inference via the introduced residual magnitude. Similarly, on a 6-station manufacturing line with constant work-in-process (CONWIP) flow control, we achieve 28.4% lower RMSE than classical LSTM baselines; a high starvation-event recall (1.0 in test cycles), and effective bottleneck identification (Precision@1 = 1.0 in test set, 13x signal-to-noise ratio). We then certify the dispatch policies with Hamilton-Jacobi reachability theory which reduces deadlock events by 96% compared to uncontrolled dispatch over 100 simulated runs (351 events prevented).

扩散模型控制策略鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。