提出跨范式注意力机制,提升雨滴图像修复的全局与局部特征融合能力。
Cross Paradigm Representation and Alignment Transformer for Image Deraining
- 设计双通道自注意力结构,融合空间与通道维度的全局感知能力。
- 在8个基准数据集上达到当前最优性能,显著改善复杂雨痕重建效果。
- 适合图像去雨、图像修复等低层视觉任务研究者参考。
基于Transformer的网络在图像去雨等低层视觉任务中表现优异,主要依赖空间或通道自注意力机制。然而,不规则雨痕和复杂的几何重叠挑战单一范式架构,亟需统一框架整合互补的全局-局部与空间-通道表征。为此,我们提出跨范式表征与对齐Transformer(CPRAformer),其核心为分层表征与对齐机制,充分利用空间-通道与全局-局部范式的协同优势,促进特征深度交互与融合。具体地,在Transformer模块中引入两种自注意力:稀疏提示通道自注意力(SPC-SA)通过动态稀疏性增强全局通道依赖,空间像素精修自注意力(SPR-SA)聚焦空间雨分布与细粒度纹理恢复。为解决范式内及范式间特征错位与知识差异,提出自适应对齐频率模块(AAFM),采用两阶段渐进式对齐策略,实现特征自适应引导与互补,缩小范式间信息差距。通过这一统一的跨范式动态交互框架,有效提取两范式间最具价值的融合信息。大量实验表明,该模型在8个基准数据集上均达领先水平,并验证了其在其他图像修复任务及下游应用中的鲁棒性。
原文摘要 · Abstract (English)
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。