提出解耦训练方法,加速非线性变换的图像压缩模型收敛。
On Disentangled Training for Nonlinear Transform in Learned Image Compression
- 引入线性辅助变换解耦能量集中问题
- 训练时间从两周以上缩短至数天
- 适合追求高效训练的图像压缩研究者
学习型图像压缩(LIC)相比传统编码器展现出更优的率失真性能,但面临训练效率低下的挑战,从头训练一个先进模型可能需要超过两周时间。现有LIC方法忽略了非线性变换中因能量集中导致的收敛缓慢问题。本文首次揭示该能量集中包含特征去相关与不均等能量调制两个成分。基于此,提出线性辅助变换(AuxT),通过粗略近似实现高效能量集中,使非线性变换的分布拟合简化为细节优化。进一步设计基于小波的线性捷径(WLSs),利用小波下采样和正交线性投影实现特征去相关,以及子带感知缩放,有效加速训练过程。
原文摘要 · Abstract (English)
Learned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, i.e., feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。