提出可保持尺度一致的扩散逆问题后验采样方法,提升超分辨率与去模糊重建质量。
Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems

- 构建可控随机性参数的后验SDE族,实现概率流与探索的解耦
- 在FFHQ和ImageNet上仅用100次评分评估即达竞争性重建效果
- 适合需高精度图像重建且关注采样稳定性研究的科研人员
使用预训练扩散先验进行后验采样的过程受制于一个通常不可计算的中间似然项。本文从一个理想的一参数后验SDE族出发,其中随机性参数可在不改变后验边际的前提下调控概率流传输与随机探索。为获得可计算模型,将似然项表达在重缩放的干净图像坐标系中,并利用log-SNR组织后验代理。通过前向算子投影扩散不确定性,得到噪声条件下的协方差路径,其目标趋近于真实后验。由于终点一致性无法保证代理传输路径与其一致,我们引入冻结目标的Langevin校正器,实现连续代理SDE。采用外层Lie-Trotter分裂与方差匹配的分裂步隐式-显式预测器离散化该模型,显式处理学习到的先验,隐式处理线性似然,最后处理随机增量。证明了理想族的边缘不变性,连续代理在混合与传输缺陷条件下对后验的收敛性,以及离散算法的一阶弱误差界。在FFHQ和ImageNet上的实验表明,仅需100次得分评估即可实现具有竞争力的超分辨率与去模糊重建。100张图像的消融实验分离出尺度一致性与有限步长下随机增量放置、延续及校正分配的影响。无噪声盒内补全研究显示,大范围探索仅在匹配的增量注入至刚性似然求解之后才能达到性能平台期。
原文摘要 · Abstract (English)
Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generally intractable. We begin from an ideal one-parameter posterior SDE family in which a stochasticity parameter controls probability-flow transport and stochastic exploration without changing the posterior marginals. To obtain a tractable model, we express the likelihood in a rescaled clean-image coordinate and use log-SNR to organize the resulting posterior proxies. Projecting the diffusion uncertainty through the forward operator then yields a noise-conditioned covariance path whose targets approach the clean posterior. Because endpoint consistency of these targets does not ensure that a surrogate transport follows them, we interleave the transport with a frozen-target Langevin corrector, producing a continuous surrogate SDE. We discretize this model with an outer Lie--Trotter splitting and a variance-matched split-step IMEX predictor that treats the learned prior explicitly, the linear likelihood implicitly, and the stochastic innovation after the implicit solve. We prove marginal invariance of the ideal family, posterior convergence of the continuous surrogate under mixing and transport-defect conditions, and a first-order weak error bound for the discrete algorithm. Experiments on FFHQ and ImageNet with 100 score evaluations demonstrate competitive reconstruction fidelity for super-resolution and deblurring. A controlled 100-image ablation separates scale consistency from the finite-step effects of stochastic-increment placement, continuation, and corrector allocation. A separate noiseless box-inpainting study shows that large exploration reaches a performance plateau only when the matched innovation is injected after the stiff likelihood solve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。