arXiv:2511.16870cs.CVcs.LG2025-11

用视觉模型对齐提升生成模型在逆问题中的重建质量。

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

  • 通过对齐扩散或流模型与DINOv2的内部表征来引导重建。
  • 在超分辨率、去模糊等任务中显著提升图像质量和效率,减少迭代步数。
  • 适用于图像恢复领域研究者,尤其关注生成模型与自监督特征对齐的场景。

将扩散或流式生成模型的内部表征与预训练自监督编码器(如DINOv2)对齐,已被证明能有效提升收敛速度与生成样本质量。本文将该思想扩展至逆问题求解,利用预训练生成模型作为先验,在推理阶段引入表示对齐(REPA)机制。尽管逆问题中缺乏真实信号,但实验证明,对近似目标特征的表征进行对齐可显著改善重建质量与感知真实感。理论分析表明:(a) REPA正则化可视为在DINOv2嵌入空间中最小化某种散度的变分方法;(b) 在特定正则性假设下,REPA更新可引导潜变量扩散状态向干净图像的潜空间靠近。这些结果揭示了REPA在提升感知保真度中的作用。我们还将REPA集成至多个先进逆问题求解器,在超分辨率、框内补全、高斯模糊去除和运动模糊去除等任务上进行了广泛实验,结果表明该方法在保持高质量的同时,大幅减少所需离散化步数,兼具性能与效率优势。

原文摘要 · Abstract (English)

Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encoders has recently been shown to provide a powerful inductive bias, improving both convergence and sample quality. In this work, we extend this idea to inverse problems, where pretrained generative models are employed as priors. We propose applying representation alignment (REPA) between diffusion or flow-based models and a DINOv2 visual encoder, to guide the reconstruction process at inference time. Although ground-truth signals are unavailable in inverse problems, we empirically show that aligning model representations of approximate target features can substantially enhance reconstruction quality and perceptual realism. We provide theoretical results showing (a) that REPA regularization can be viewed as a variational approach for minimizing a divergence measure in the DINOv2 embedding space, and (b) how under certain regularity assumptions REPA updates steer the latent diffusion states toward those of the clean image. These results offer insights into the role of REPA in improving perceptual fidelity. Finally, we demonstrate the generality of our approach by We integrate REPA into multiple state-of-the-art inverse problem solvers, and provide extensive experiments on super-resolution, box inpainting, Gaussian deblurring, and motion deblurring confirming that our method consistently improves reconstruction quality, while also providing efficiency gains reducing the number of required discretization steps.

图像重建扩散模型表征对齐逆问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。