arXiv:2506.18437cs.CV2025-06被引 2

用频域融合提升Transformer图像修复细节保留能力

Frequency-Domain Fusion Transformer for Image Inpainting

  • 结合小波与吉布斯滤波的注意力机制,增强多尺度结构建模
  • 通过可学习频域滤波器,自适应抑制噪声并保留高频细节
  • 适合需要高精度细节修复的图像重建任务

图像修复在恢复缺失区域和支撑高层视觉任务中至关重要,但传统方法在复杂纹理和大范围遮挡下表现不佳。尽管基于Transformer的方法具备强大的全局建模能力,却因自注意力的低通特性难以保留高频细节,且计算成本较高。为此,本文提出一种结合频域融合的Transformer图像修复方法。具体地,引入结合小波变换与吉布斯滤波的注意力机制,以增强多尺度结构建模与细节保持能力;同时设计基于快速傅里叶变换的可学习频域滤波器,替代传统前馈网络,实现自适应噪声抑制与细节保留。模型采用四级编码器-解码器结构,并通过新型损失策略平衡全局语义与精细细节。实验表明,该方法能有效提升图像修复质量,显著保留更多高频信息。

原文摘要 · Abstract (English)

Image inpainting plays a vital role in restoring missing image regions and supporting high-level vision tasks, but traditional methods struggle with complex textures and large occlusions. Although Transformer-based approaches have demonstrated strong global modeling capabilities, they often fail to preserve high-frequency details due to the low-pass nature of self-attention and suffer from high computational costs. To address these challenges, this paper proposes a Transformer-based image inpainting method incorporating frequency-domain fusion. Specifically, an attention mechanism combining wavelet transform and Gabor filtering is introduced to enhance multi-scale structural modeling and detail preservation. Additionally, a learnable frequency-domain filter based on the fast Fourier transform is designed to replace the feedforward network, enabling adaptive noise suppression and detail retention. The model adopts a four-level encoder-decoder structure and is guided by a novel loss strategy to balance global semantics and fine details. Experimental results demonstrate that the proposed method effectively improves the quality of image inpainting by preserving more high-frequency information.

图像修复Transformer频域融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。