揭示扩散模型采样误差如何随数据几何变化,指导参数优化
Geometry-Aware Discretization Error of Diffusion Models

- 基于欧拉-马鲁亚玛方法推导误差的渐近展开式
- 误差受数据协方差谱与扩散参数共同影响,可显式表达
- 公式在不同数据集上保持稳健,适用于图像生成与后验采样
实际扩散采样本质是数值近似问题:在固定推理预算下,需用有限步数模拟反向时间常微分方程或随机微分方程,因此离散化误差常成为主要误差来源。现有非渐近分析虽提供收敛保证,但通常过松且对扩散参数不敏感,导致不同调度方案获得相同误差率,依赖如维度或漂移Lipschitz常数等粗粒度最坏情况量。本文采取更务实但更具信息量的路径:在精确得分设定下,推导欧拉-马鲁亚玛方法的弱误差与Fréchet离散化误差的一阶渐近展开式。该公式适用于一般光滑反向扩散过程,在高斯数据下可完全显式表达。结果表明,离散化误差随数据协方差谱演变,且与扩散调度、扩散项系数等关键参数交互作用。由此可构建可计算的目标函数,实现几何感知的参数优化。最后,我们验证高斯公式中的定性预测在不同几何结构的问题中仍具鲁棒性,包括不同数据集上的图像生成与图像后验采样。
原文摘要 · Abstract (English)
Practical diffusion sampling is a numerical approximation problem: under a fixed inference budget, one must simulate a reverse-time ODE or SDE using only a limited number of denoising steps, so discretization error is often the dominant source of error. Existing non-asymptotic analyses provide convergence guarantees, but are typically too loose and too insensitive to diffusion parameters to guide practical design: broad families of schedules receive the same rates, which depend on coarse worst-case quantities such as the dimension or the drift Lipschitz constant. We take a less ambitious but more informative route. In the exact-score setting, we derive first-order asymptotic expansions of the Euler-Maruyama weak and Fréchet discretization errors. These formulas hold for general smooth reverse diffusions and become fully explicit under Gaussian data. They show how discretization error adapts to the geometry of the data through the covariance spectrum, and how this geometry interacts with key diffusion parameters, including the diffusion schedules and the diffusion-term coefficient. This yields tractable objectives for geometry-aware parameter optimization. Finally, we show that the qualitative predictions of the Gaussian formulas remain robust across diffusion sampling problems with different geometries, including image generation on different datasets and image posterior sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。