一个模型搞定所有图像修复,靠视觉指令精准识别退化类型
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- 用标准图像加退化生成视觉指令,明确对应各类图像退化特征
- 在退化空间中通过扩散模型去噪,实现稳定且泛化性强的修复
- 适合处理真实场景中混合或未知退化的复杂图像修复任务
图像修复任务如去模糊、去噪、去雾通常需要针对每种退化类型设计独立模型,限制了在实际场景中混合或未知退化下的泛化能力。本文提出全新的统一图像修复框架Defusion,采用视觉指令引导的退化扩散机制。与依赖特定任务模型或模糊文本先验的方法不同,Defusion通过在标准化视觉元素上施加退化,构建与视觉退化模式对齐的显式视觉指令,捕捉内在退化特征且不依赖图像语义。随后,这些视觉指令引导基于扩散的模型直接在退化空间中运行,通过增强稳定性与泛化性,实现高质量图像重建。大量实验表明,Defusion在多种图像修复任务中均超越现有最先进方法,包括复杂及真实世界退化场景。
原文摘要 · Abstract (English)
Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose \textbf{Defusion}, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit \textbf{visual instructions} that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。