arXiv:2505.18047cs.CVcs.AI2025-05被引 13

用自回归模型提升图像修复速度与效果,10倍快于主流扩散模型。

RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration

  • 采用视觉自回归架构,分尺度逐步生成图像。
  • 修复精度超越扩散模型,推理速度提升10倍以上。
  • 适合需要快速高质图像修复的场景,如实时视频处理。

基于潜在扩散模型(LDM)如Stable Diffusion的全模态图像修复(AiOR)方法在感知质量与泛化能力上取得显著进展,但其迭代去噪过程导致推理缓慢,难以满足时间敏感应用需求。视觉自回归建模(VAR)是一种新兴的图像生成方法,通过尺度空间自回归实现与顶尖扩散变压器相当的性能,且计算成本大幅降低。我们的分析表明,VAR的粗尺度主要捕捉退化特征,细尺度则编码场景细节,这为修复任务提供了清晰的层次结构。受此启发,我们提出RestoreVAR——一种基于VAR的新型生成式全模态图像修复方法,在修复性能上显著优于现有LDM模型,同时实现超过10倍的推理加速。为充分发挥VAR在AiOR任务中的优势,我们设计了定制化的交叉注意力机制和潜空间精修模块。大量实验表明,RestoreVAR在生成式AiOR方法中达到最先进水平,并具备强大的泛化能力。

原文摘要 · Abstract (English)

The use of latent diffusion models (LDMs) such as Stable Diffusion has significantly improved the perceptual quality of All-in-One image Restoration (AiOR) methods, while also enhancing their generalization capabilities. However, these LDM-based frameworks suffer from slow inference due to their iterative denoising process, rendering them impractical for time-sensitive applications. Visual autoregressive modeling (VAR), a recently introduced approach for image generation, performs scale-space autoregression and achieves comparable performance to that of state-of-the-art diffusion transformers with drastically reduced computational costs. Moreover, our analysis reveals that coarse scales in VAR primarily capture degradations while finer scales encode scene detail, simplifying the restoration process. Motivated by this, we propose RestoreVAR, a novel VAR-based generative approach for AiOR that significantly outperforms LDM-based models in restoration performance while achieving over $10\times$ faster inference. To optimally exploit the advantages of VAR for AiOR, we propose architectural modifications and improvements, including intricately designed cross-attention mechanisms and a latent-space refinement module, tailored for the AiOR task. Extensive experiments show that RestoreVAR achieves state-of-the-art performance among generative AiOR methods, while also exhibiting strong generalization capabilities.

图像修复自回归生成模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。