arXiv:2503.14037cs.CV2025-03AAAI被引 14

用解析提示增强视觉模型,提升图像修复效果

Intra and Inter Parser-Prompted Transformers for Effective Image Restoration

论文配图:Intra and Inter Parser-Prompted Transformers for Effective Image Restoration
图 1 · 摘自论文原文
  • 引入解析提示注意力机制,隐式与显式融合解析信息
  • 在去雨、模糊、降雪、低光修复上达到顶尖性能
  • 适合图像修复研究者和实际应用开发者参考

我们提出一种基于解析提示的变压器架构(PPTformer),用于从视觉基础模型中挖掘有用特征以实现图像修复。PPTformer包含两部分:用于从退化观测中恢复图像的图像修复网络(IRNet)和为IRNet提供可靠解析信息以增强修复效果的解析提示特征生成网络(PPFGNet)。为加强解析信息与IRNet的融合,我们设计了内部解析提示注意力(IntraPPA)和跨层解析提示注意力(InterPPA),分别从层内长程视角隐式感知解析特征,以及通过特征融合与注意力机制显式感知解析信息。此外,还提出解析提示前馈网络,通过像素级门控调制引导修复过程。实验表明,PPTformer在去雨、散焦去模糊、去雪和低光增强任务上均取得当前最优性能。

原文摘要 · Abstract (English)

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted Feature Generation Network (PPFGNet) for providing IRNet with reliable parser information to boost restoration. To enhance the integration of the parser within IRNet, we propose Intra Parser-Prompted Attention (IntraPPA) and Inter Parser-Prompted Attention (InterPPA) to implicitly and explicitly learn useful parser features to facilitate restoration. The IntraPPA re-considers cross attention between parser and restoration features, enabling implicit perception of the parser from a long-range and intra-layer perspective. Conversely, the InterPPA initially fuses restoration features with those of the parser, followed by formulating these fused features within an attention mechanism to explicitly perceive parser information. Further, we propose a parser-prompted feed-forward network to guide restoration within pixel-wise gating modulation. Experimental results show that PPTformer achieves state-of-the-art performance on image deraining, defocus deblurring, desnowing, and low-light enhancement.

图像修复注意力机制视觉模型Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。