通过3D多级小波卷积提升图像压缩率与质量
3DM-WeConvene: Learned Image Compression with 3D Multi-Level Wavelet-Domain Convolution and Entropy Model
- 将小波变换嵌入3D卷积层,分频段处理不同尺度特征
- 在高分辨率图像上相比H.266/VVC降低12.97%以上码率
- 适合追求高压缩效率的图像编码研究与应用
学习型图像压缩(LIC)近年来取得显著进展,超越传统方法。然而,多数LIC方法主要在空间域操作,缺乏对频域相关性的建模能力。为此,我们提出一种新框架,将低复杂度的3D多级离散小波变换(DWT)集成到卷积层与熵编码中,同时减少空间与通道相关性,提升频域选择性与率失真(R-D)性能。提出的3D多级小波域卷积(3DM-WeConv)层首先对数据应用3D多级DWT(如JPEG 2000中的5/3和9/7小波),将其转换至小波域;随后对不同频带应用不同尺寸卷积,再经逆3D DWT恢复空间域。该层可灵活嵌入现有基于CNN的LIC模型。我们还引入3D小波域通道自回归熵模型(3DWeChARM),在3D DWT域进行分片熵编码,优先编码低频(LF)片以提供高频(HF)片的先验信息。采用两阶段训练策略:先平衡高低频码率,再分别加权微调。大量实验表明,本框架在率失真性能与计算复杂度上持续优于当前最先进的基于CNN的LIC方法,尤其在高分辨率图像上增益更显著。在Kodak、Tecnick 100和CLIC测试集上,相比H.266/VVC分别实现-12.24%、-15.51%和-12.97%的BD-Rate降低。
原文摘要 · Abstract (English)
Learned image compression (LIC) has recently made significant progress, surpassing traditional methods. However, most LIC approaches operate mainly in the spatial domain and lack mechanisms for reducing frequency-domain correlations. To address this, we propose a novel framework that integrates low-complexity 3D multi-level Discrete Wavelet Transform (DWT) into convolutional layers and entropy coding, reducing both spatial and channel correlations to improve frequency selectivity and rate-distortion (R-D) performance. Our proposed 3D multi-level wavelet-domain convolution (3DM-WeConv) layer first applies 3D multi-level DWT (e.g., 5/3 and 9/7 wavelets from JPEG 2000) to transform data into the wavelet domain. Then, different-sized convolutions are applied to different frequency subbands, followed by inverse 3D DWT to restore the spatial domain. The 3DM-WeConv layer can be flexibly used within existing CNN-based LIC models. We also introduce a 3D wavelet-domain channel-wise autoregressive entropy model (3DWeChARM), which performs slice-based entropy coding in the 3D DWT domain. Low-frequency (LF) slices are encoded first to provide priors for high-frequency (HF) slices. A two-step training strategy is adopted: first balancing LF and HF rates, then fine-tuning with separate weights. Extensive experiments demonstrate that our framework consistently outperforms state-of-the-art CNN-based LIC methods in R-D performance and computational complexity, with larger gains for high-resolution images. On the Kodak, Tecnick 100, and CLIC test sets, our method achieves BD-Rate reductions of -12.24%, -15.51%, and -12.97%, respectively, compared to H.266/VVC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。