arXiv:2606.00111eess.IVcs.CV2026-06

用小波域通道注意力提升图像压缩率,性能显著优于现有方法。

ChWDTA: Channel-wise Wavelet-Domain Transformer Attention and Entropy Modeling for Learned Image Compression

论文配图:ChWDTA: Channel-wise Wavelet-Domain Transformer Attention and Entropy Modeling for Learned Image Compression
图 1 · 摘自论文原文
  • 在通道级小波域进行注意力计算,保留窗口化自注意机制
  • 在Kodak等数据集上实现最高达22.56%的码率降低
  • 适合追求高效率图像压缩的研究与工程人员

当前最先进的学习型图像压缩(LIC)方案越来越多地采用混合卷积神经网络-变压器架构。为进一步提升率失真性能,本文在变压器和熵编码组件中引入通道级小波变换。首先提出通道级小波域变压器注意力(ChWDTA)机制:在保持现代LIC主干中高效窗口化空间自注意力的基础上,将查询/键/值投影在通道级小波变换特征上,再通过逆变换映射回原域。由此产生的通道级小波域变压器块(ChWDTB)在保留窗口注意力空间标记模式的同时,稀疏化了注意力投影所见的通道协方差。其次,在熵编码阶段,引入通道级小波包分解(ChWP),生成四个等大小子带,更适配通道级切片式自回归熵建模。当每个通道级子带被分为两个切片时,共使用八个切片进行熵编码。在此配置下,该方案在Kodak、CLIC专业验证集和Tecnick测试集上分别取得-17.82%、-19.15%和-22.56%的BD-rate降低。即使每个通道级子带作为单一切片编码,仍保持大部分编码增益且复杂度更低。结果验证了在基于CNN-Transformer的LIC方案中引入小波变换的优势。

原文摘要 · Abstract (English)

State-of-the-art learned image compression (LIC) schemes are increasingly based on hybrid CNN-transformer architectures. To further improve rate-distortion performance, we introduce channel-wise wavelet transforms into both the transformer and entropy-coding components. First, we propose a channel-wise wavelet-domain transformer attention (ChWDTA) mechanism. ChWDTA keeps the efficient windowed spatial self-attention used in modern LIC backbones, but computes the Q/K/V projections on channel-wise wavelet-transformed features before mapping the attention output back with the inverse transform. The resulting Channel-wise Wavelet-Domain Transformer Block (ChWDTB) therefore preserves the spatial tokenization pattern of windowed attention while sparsifying the channel covariance seen by the attention projections. Second, in the entropy-coding stage, we introduce a channel-wise wavelet packet (ChWP) decomposition that produces four equal-sized subbands, which better fit channel-wise slice-based autoregressive entropy modeling. When each channel-wise subband is divided into two slices, we use eight slices for entropy coding. With this configuration, the proposed scheme obtains BD-rate reductions of -17.82%, -19.15%, and -22.56% on the Kodak, CLIC Professional Validation, and Tecnick test sets, respectively. Even when each channel-wise subband is coded as a single slice, the scheme still retains most of the coding gains with lower complexity. The results confirm the advantage of introducing wavelet transform in CNN-transformer-based LIC schemes.

图像压缩小波变换变压器熵建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。