证明扩散采样中前向误差小不保证数值稳定,提出投影可抑制奇异轨迹。
Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling

- 用光滑得分场构造前向误差极小但反向路径发散的反例
- 即使路径空间收敛,所有Wasserstein距离仍发散至无穷
- 对固定网络架构,投影得分器可实现稳定采样,适合小模型应用
得分匹配控制前向边缘分布的平均误差,但离散化的反向时间采样器沿自身轨迹评估学习到的得分。我们证明:前向边际误差小并不保证数值稳定性。构造了一个光滑得分场,其前向边际 $L^2$ 误差任意小;学习到的反向过程非爆炸性,各阶矩存在,且在路径空间总变差意义下可任意接近精确反向过程。然而,其 Euler--Maruyama 离散化在概率意义下收敛,而所有正阶矩均发散。因此,弱收敛成立时,所有 Wasserstein 距离 $W_p$($p o1$)仍发散。该现象可在单一固定神经架构内发生。我们构造了一族有界、全局 Lipschitz 去噪器,其前向误差与路径空间总变差距离均趋于零,但 Euler--Maruyama 终点在所有 $W_p$ 意义下发散。对于紧支撑数据,给出一个正结果:将学习到的去噪器投影到包含支持的已知有界闭凸集,可保持逐点精度,获得网格一致的矩界,并在局部弱正则条件下实现 Wasserstein 收敛。实验使用小型 DiT 架构显示罕见轨迹存在剧烈增长,而投影能有效抑制此现象,整体轨迹误差仍较小。
原文摘要 · Abstract (English)
Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal $L^2$ error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance $W_p$, $p\ge1$, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every $W_p$. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。