用自回归生成模型的分布对齐能力,提升图像修复的通用性与效率。
Navigating Image Restoration with VAR's Distribution Alignment Prior
- 利用VAR模型的多尺度自回归过程,自动对齐不同层级特征分布。
- 在多种未见修复任务上实现领先性能,训练成本更低。
- 适合需要高泛化能力的图像修复场景,如低质量图像恢复。
基于大规模高质量数据集训练的生成模型能有效捕捉干净图像的结构与统计特性,成为图像修复中从退化特征恢复清晰图像的强大先验。VAR是一种新型图像生成范式,通过逐尺度预测方法,在生成质量上超越扩散模型。其自回归过程逐步捕获全局结构与细粒度细节,符合修复领域普遍认可的多尺度原则。我们观察到,在使用VAR进行图像重建时,尺度预测会自动调制输入,促进后续尺度表示与干净图像分布的对齐。为此,我们将VAR中的多尺度潜在表示作为修复先验,构建了精心设计的VarFormer框架。该先验的策略性应用使VarFormer在未见过的任务上表现出卓越泛化能力,同时降低训练计算开销。大量实验表明,我们的VarFormer在多种修复任务中均优于现有方法。
原文摘要 · Abstract (English)
Generative models trained on extensive high-quality datasets effectively capture the structural and statistical properties of clean images, rendering them powerful priors for transforming degraded features into clean ones in image restoration. VAR, a novel image generative paradigm, surpasses diffusion models in generation quality by applying a next-scale prediction approach. It progressively captures both global structures and fine-grained details through the autoregressive process, consistent with the multi-scale restoration principle widely acknowledged in the restoration community. Furthermore, we observe that during the image reconstruction process utilizing VAR, scale predictions automatically modulate the input, facilitating the alignment of representations at subsequent scales with the distribution of clean images. To harness VAR's adaptive distribution alignment capability in image restoration tasks, we formulate the multi-scale latent representations within VAR as the restoration prior, thus advancing our delicately designed VarFormer framework. The strategic application of these priors enables our VarFormer to achieve remarkable generalization on unseen tasks while also reducing training computational costs. Extensive experiments underscores that our VarFormer outperforms existing multi-task image restoration methods across various restoration tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。