将视觉变压器与频域模块结合,提升复杂模糊图像的恢复效果。
From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
- 融合空间注意力与频域稀疏性,双域协同建模图像模糊特征。
- 在多个基准数据集上超越现有模型,PSNR和SSIM均显著提升。
- 适合需要高精度图像修复的工业场景或真实世界应用。
图像去模糊在计算机视觉中至关重要,旨在从运动或相机抖动导致的模糊图像中恢复清晰图像。尽管深度学习方法如卷积神经网络(CNNs)和视觉变压器(ViTs)已推动该领域发展,但它们在处理复杂或高分辨率模糊时仍面临挑战,且计算开销较大。本文提出一种新型双域架构,将视觉变压器与频域的FFT-ReLU模块有机结合,显式连接空间注意力建模与频域稀疏性。该结构中,ViT骨干网络捕捉局部与全局依赖关系,而FFT-ReLU组件通过强制频域稀疏性,抑制模糊相关伪影并保留细节。大量实验在多个基准数据集上表明,该方法在PSNR、SSIM及感知质量方面均优于当前最优模型。定量指标、定性对比及人类偏好评估均证实其有效性,建立了一种实用且可推广的真实世界图像恢复范式。
原文摘要 · Abstract (English)
Image deblurring is vital in computer vision, aiming to recover sharp images from blurry ones caused by motion or camera shake. While deep learning approaches such as CNNs and Vision Transformers (ViTs) have advanced this field, they often struggle with complex or high-resolution blur and computational demands. We propose a new dual-domain architecture that unifies Vision Transformers with a frequency-domain FFT-ReLU module, explicitly bridging spatial attention modeling and frequency sparsity. In this structure, the ViT backbone captures local and global dependencies, while the FFT-ReLU component enforces frequency-domain sparsity to suppress blur-related artifacts and preserve fine details. Extensive experiments on benchmark datasets demonstrate that this architecture achieves superior PSNR, SSIM, and perceptual quality compared to state-of-the-art models. Both quantitative metrics, qualitative comparisons, and human preference evaluations confirm its effectiveness, establishing a practical and generalizable paradigm for real-world image restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。