arXiv:2411.18003eess.IVcs.AI2024-11被引 2

HAAT通过混合注意力机制提升图像超分辨率质量

HAAT: Hybrid Attention Aggregation Transformer for Image Super-Resolution

  • 融合窗口、稀疏与通道注意力,增强跨域特征提取
  • 在DIV2K和Urban100上峰值信噪比分别达31.89和29.72
  • 适合追求高细节图像重建的视觉算法研究者

在图像超分辨率领域,基于Swin-Transformer的模型因具备全局空间建模能力和移位窗口注意力机制而备受青睐。然而,现有方法常将自注意力限制在非重叠窗口以降低成本,忽略了跨通道的有用信息。为此,本文提出一种新型模型——混合注意力聚合变换器(HAAT),旨在更充分地利用特征信息。HAAT由Swin-密集残差连接块(SDRCB)与混合网格注意力块(HGAB)构成。SDRCB在保持轻量架构的同时扩展感受野,提升性能;HGAB融合通道注意力、稀疏注意力与窗口注意力,增强非局部特征融合,实现更逼真的视觉效果。实验表明,HAAT在多个基准数据集上超越现有最先进方法。

原文摘要 · Abstract (English)

In the research area of image super-resolution, Swin-transformer-based models are favored for their global spatial modeling and shifting window attention mechanism. However, existing methods often limit self-attention to non overlapping windows to cut costs and ignore the useful information that exists across channels. To address this issue, this paper introduces a novel model, the Hybrid Attention Aggregation Transformer (HAAT), designed to better leverage feature information. HAAT is constructed by integrating Swin-Dense-Residual-Connected Blocks (SDRCB) with Hybrid Grid Attention Blocks (HGAB). SDRCB expands the receptive field while maintaining a streamlined architecture, resulting in enhanced performance. HGAB incorporates channel attention, sparse attention, and window attention to improve nonlocal feature fusion and achieve more visually compelling results. Experimental evaluations demonstrate that HAAT surpasses state-of-the-art methods on benchmark datasets. Keywords: Image super-resolution, Computer vision, Attention mechanism, Transformer

图像超分辨率Transformer注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。