arXiv:2412.00878cs.CV2024-12被引 17

用文本增强提升图像修复模型在真实场景下的泛化能力

Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration

  • 引入文本作为辅助不变表示,激活扩散模型的生成能力
  • 新模块Res-Captioner根据图像内容生成优化文本描述,提升鲁棒性
  • 适用于需要强泛化能力的真实图像修复任务

真实世界图像修复的泛化能力长期是核心挑战。尽管基于扩散的修复方法借助文本到图像模型的生成先验,在恢复更真实细节方面取得进展,但在分布外的真实数据上仍会出现“生成能力失效”现象。为此,我们提出利用文本作为辅助不变表示,重新激活模型的生成能力。通过分析文本输入的丰富性与相关性两个关键属性对性能的影响,我们设计了Res-Captioner模块,可生成针对图像内容与退化程度定制的增强文本描述,有效缓解响应失败问题。此外,我们构建了RealIR新基准,以涵盖多样化的现实场景。大量实验表明,Res-Captioner显著提升了扩散模型的泛化能力,且完全即插即用。

原文摘要 · Abstract (English)

Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-image models, have made progress in recovering more realistic details, they still encounter "generative capability deactivation" when applied to out-of-distribution real-world data. To address this, we propose using text as an auxiliary invariant representation to reactivate the generative capabilities of these models. We begin by identifying two key properties of text input: richness and relevance, and examine their respective influence on model performance. Building on these insights, we introduce Res-Captioner, a module that generates enhanced textual descriptions tailored to image content and degradation levels, effectively mitigating response failures. Additionally, we present RealIR, a new benchmark designed to capture diverse real-world scenarios. Extensive experiments demonstrate that Res-Captioner significantly enhances the generalization abilities of diffusion-based restoration models, while remaining fully plug-and-play.

图像修复扩散模型文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。