用双提示提升扩散模型修复质量
Dual Prompting Image Restoration with Diffusion Transformers
- 双分支设计:图像先验与全局局部视觉提示并行
- 在Real-World Deblur等数据集上优于现有方法
- 适合追求高保真图像修复的研究者
近期主流图像修复方法多采用基于U-Net的潜在扩散模型,但受限于能力瓶颈,仍难以实现高质量修复。扩散 Transformer(DiT)如SD3因其可扩展性与更优画质成为新兴替代方案。本文提出DPIR(Dual Prompting Image Restoration),通过多视角提取低质图像的条件信息。DPIR包含两个分支:低质量图像条件分支与双提示控制分支。前者采用轻量模块高效融入图像先验;更重要的是,我们认为仅靠文本描述无法充分捕捉图像丰富的视觉特征,因此设计双提示模块,提供全局上下文与局部外观双重视觉线索。提取的全局-局部视觉提示作为额外条件控制,结合文本提示形成双提示,显著提升修复质量。大量实验表明,DPIR在Real-World Deblur、GoPro等数据集上均取得领先性能。
原文摘要 · Abstract (English)
Recent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because of their better quality with scalability. In this paper, we introduce DPIR (Dual Prompting Image Restoration), a novel image restoration method that effectivly extracts conditional information of low-quality images from multiple perspectives. Specifically, DPIR consits of two branches: a low-quality image conditioning branch and a dual prompting control branch. The first branch utilizes a lightweight module to incorporate image priors into the DiT with high efficiency. More importantly, we believe that in image restoration, textual description alone cannot fully capture its rich visual characteristics. Therefore, a dual prompting module is designed to provide DiT with additional visual cues, capturing both global context and local appearance. The extracted global-local visual prompts as extra conditional control, alongside textual prompts to form dual prompts, greatly enhance the quality of the restoration. Extensive experimental results demonstrate that DPIR delivers superior image restoration performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。