通过窗口化通道注意力提升图像压缩的全局感知能力
Window-based Channel Attention for Wavelet-enhanced Learned Image Compression
- 在通道注意力中引入窗口划分,扩展感受野
- 结合离散小波变换实现频域下采样,进一步扩大感知范围
- 在四个数据集上比VTM-23.1降低20%以上BD-rate,适合高保真压缩场景
学习型图像压缩(LIC)模型在率失真性能上已超越传统编码器。现有LIC模型多采用CNN、Transformer或混合结构。然而,基于Swin-Transformer的LIC受限于移位窗口注意力,感受野增长受限,影响对大物体的建模能力。为此,本文首次将窗口划分引入通道注意力,以获得更大的感受野并捕捉更多全局信息。由于通道注意力会抑制局部信息学习,有必要将现有Transformer编码器中的注意力机制扩展为空间-通道混合注意力,建立多重感受野,既能捕获大范围相关性,又能保持小尺度局部细节。此外,我们在空间-通道混合(SCH)框架中引入离散小波变换,实现高效的频率依赖下采样,进一步扩大感受野。实验表明,该方法在四个标准数据集上相比VTM-23.1分别降低了18.54%、23.98%、22.33%和24.71%的BD-rate,达到当前最优性能。
原文摘要 · Abstract (English)
Learned Image Compression (LIC) models have achieved superior rate-distortion performance than traditional codecs. Existing LIC models use CNN, Transformer, or Mixed CNN-Transformer as basic blocks. However, limited by the shifted window attention, Swin-Transformer-based LIC exhibits a restricted growth of receptive fields, affecting the ability to model large objects for image compression. To address this issue and improve the performance, we incorporate window partition into channel attention for the first time to obtain large receptive fields and capture more global information. Since channel attention hinders local information learning, it is important to extend existing attention mechanisms in Transformer codecs to the space-channel attention to establish multiple receptive fields, being able to capture global correlations with large receptive fields while maintaining detailed characterization of local correlations with small receptive fields. We also incorporate the discrete wavelet transform into our Spatial-Channel Hybrid (SCH) framework for efficient frequency-dependent down-sampling and further enlarging receptive fields. Experiment results demonstrate that our method achieves state-of-the-art performances, reducing BD-rate by 18.54%, 23.98%, 22.33%, and 24.71% on four standard datasets compared to VTM-23.1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。