arXiv:2608.23916cs.LGcond-mat.stat-mech2026-08

揭示扩散模型训练损失的理论下限及其信息几何根源

The Loss Floor of Denoising Score Matching: Fisher Geometry from Schrödinger Bridges

  • 从薛定谔桥视角推导出条件方差分解,定位损失下限来源
  • 损失下限等于路径上条件分布的Fisher-Rao迹积分,与信息流失率相关
  • 解释不同噪声设置下损失不可比,提示评估需考虑底层几何

去噪得分匹配通过回归条件得分来训练扩散模型,但生成最终依赖边际得分。两者在总体最小值处一致,但在固定噪声状态下,条件目标仍具随机性,导致训练损失中存在不可消除的超额项。本文将该超额项分离,并证明在温和正则性假设下,其精确等于沿扩散轨迹的条件终点族的Fisher-Rao度量迹的积分。这给出了去噪目标的精确条件方差分解,揭示扩散潜空间中观察到的信息几何是训练损失的内在组成部分。结果由薛定谔桥变分原理导出,理想目标表现为路径空间上的相对熵超额。对于污染扩散过程,该Fisher项正比于噪声状态丢失关于原始数据互信息的速率,将损失下限分为由数据决定的信息流和由污染调度及目标决定的权重。在高斯情形下,得到下限的闭式表达,恢复连续时间目标的重参数化不变性,并将高信噪比发散关联至数据的信息维数。最后,我们指出不同噪声范围或加权下的原始损失未必能一致排名模型,因其包含不同的加性下限;并对比训练所见的二阶几何与采样误差中涉及的三阶条件统计。

原文摘要 · Abstract (English)

Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher--Rao metric of the conditional endpoint family, integrated along the diffusion trajectory. This gives an exact conditional-variance decomposition of the denoising objective and identifies the information geometry observed in diffusion latent spaces as an intrinsic component of the training loss. We derive the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and objective. In the Gaussian case, this yields a closed form for the floor, recovers reparametrization invariance of the continuous-time objective, and relates its high-SNR divergence to the information dimension of the data. Finally, we show that raw losses obtained with different noise ranges or weightings need not rank models consistently because they contain different additive floors, and contrast the second-order geometry seen by training with the third-order conditional statistics entering numerical sampling error.

扩散模型信息几何损失分析薛定谔桥

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。