一次扩散反演即为凸优化,揭示模型几何特性与误差边界。
One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion

- 将扩散反演一步建模为凸势函数的驻点问题,无需假设数据流形结构。
- 在贝叶斯极限下解唯一,模型误差可由后验协方差界直接验证。
- 提出无模型依赖的收敛域判据,适合研究扩散模型内部几何与训练缺陷者。
一次隐式DDIM反演是探测预训练扩散模型是否编码局部流形几何的最廉价手段。它对应显式势函数 $x-G(x)=\nablaΨ_t(x)$ 的驻点条件,在贝叶斯极限下该势函数严格凸,其凸性模为 $e^{-h_t}$,其中 $h_t$ 为步长的对数信噪比差,适用于任意数据分布、调度和点,无需流形、可达性或单峰性假设。三个结论需区分:(i) 贝叶斯极限下解唯一;若存在第二解,则训练得分违反后验协方差界 $1/(1-e^{-h_t})$,构成无假设的模型误差证;同一界保证收缩率恒定 $ρ_g^{∗}=1-e^{-h_t}<0.326$,贯穿标准DDPM调度。(ii) 求解器仍可能失败:皮卡德迭代等价于 $Ψ_t$ 上的单位步长梯度下降,当 $λ_{\max}(\nabla^2Ψ_t)>2$ 时不稳定,振荡不表示无效;将步长减至 $2/λ_{\max}$ 可修复。(iii) 几何信息存在于收敛域中:在无量纲深度 $w=rκ_{\max}$ 下,振荡壳位于 $w=\tfrac12$(与调度无关),发散壳在 $w=1/(1+ρ_g^{∗})$,且测得有限噪声修正项 $\|\mathrm{II}\|^2$。精确得分在三类数据上复现二者误差低于 $0.54\%$;所有探测的训练得分均未显示壳结构——非空结果,而是推导出的局限:费米窗口与模型自身训练支持冲突达 $3.6$-$5.6\times$,训练海森-利普希茨常数仅为真实曲率的 $2$-$12\%$,在ReLU网络上为 $0$。最后,由 $\mathrm{Cov}(x_0\mid x_t)\succeq0$ 导出的无条件上限 $σ_tλ_{\max}(\mathrm{sym}\,J)\le1$ 对精确得分成立至 $3\times10^{-7}$,但在所有DDPM CIFAR-10/CelebA-HQ-256设置中被违反,幅度 $1.26$-$4.66\times$。
原文摘要 · Abstract (English)
One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nablaΨ_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law, schedule and point, with no manifold, reach or unimodality hypothesis. Three consequences must be kept apart. (i) The solution is unique at the Bayes limit; a second one requires the trained score to violate the posterior-covariance bound by $1/(1-e^{-h_t})$, a hypothesis-free certificate of model error; the same bound makes contraction a schedule constant, $ρ_g^{\star}=1-e^{-h_t}<0.326$ throughout the standard DDPM schedule. (ii) The solver can still fail: Picard iteration is unit-step gradient descent on $Ψ_t$, unstable wherever $λ_{\max}(\nabla^2Ψ_t)>2$, so oscillation certifies nothing; damping below $2/λ_{\max}$ cures it. (iii) The geometry lives in the convergence domain: on the scale-free depth $w=rκ_{\max}$ the oscillation shell sits at $w=\tfrac12$, schedule-free, and the divergence shell at $w=1/(1+ρ_g^{\star})$, with a measured finite-noise correction in $\|\mathrm{II}\|^2$. Exact scores reproduce both to within $0.54\%$ on three classes; no trained score we probe shows a shell $-$ a derived limitation, not a null result: the Fermi window conflicts with the model's own training support by $3.6$-$5.6\times$, and the trained Hessian-Lipschitz constant is $2$-$12\%$ of the curvature the law reads, $0$ on a ReLU net. Finally the unconditional ceiling $σ_tλ_{\max}(\mathrm{sym}\,J)\le1$, from $\mathrm{Cov}(x_0\mid x_t)\succeq0$ alone, holds for the exact score to $3\times10^{-7}$ but is violated in all DDPM CIFAR-10/CelebA-HQ-256 settings, by $1.26$-$4.66\times$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。