arXiv:2502.15174eess.IVcs.CV2025-02被引 6

针对屏幕图像压缩难题,提出多频特征分解新方法,显著提升画质。

FD-LSCIC: Frequency Decomposition-based Learned Screen Content Image Compression

  • 分频域提取特征,融合多尺度信息,减少频率间冗余
  • 自适应量化模块按频段调节精度,提升压缩灵活性
  • 构建超万张屏幕图像数据集,推动学习式压缩发展

学习型图像压缩(LIC)在自然场景图像上已超越传统方法,但直接用于具有锐边、重复纹理、文字与图形等特性的屏幕内容(SC)图像时效果不佳。本文针对学习紧凑特征表示、自适应量化步长和缺乏大规模SC数据集三大挑战,提出一种基于多频分解的压缩方法:采用多频两阶段八度残差块(MToRB)提取特征,级联三尺度特征融合残差块(CTSFRB)整合多尺度信息,以及多频上下文交互模块(MFCIM)降低频域相关性;设计自适应量化模块,为各频段学习可缩放的均匀噪声,实现量化粒度灵活控制;并构建包含超10,000张图像的SDU-SCICD10K数据集,覆盖桌面与移动端的基本屏幕图像、渲染图像及自然与屏幕混合图像。实验表明,该方法在峰值信噪比(PSNR)与多尺度结构相似性(MS-SSIM)上均优于传统标准与现有先进学习方法。

原文摘要 · Abstract (English)

The learned image compression (LIC) methods have already surpassed traditional techniques in compressing natural scene (NS) images. However, directly applying these methods to screen content (SC) images, which possess distinct characteristics such as sharp edges, repetitive patterns, embedded text and graphics, yields suboptimal results. This paper addresses three key challenges in SC image compression: learning compact latent features, adapting quantization step sizes, and the lack of large SC datasets. To overcome these challenges, we propose a novel compression method that employs a multi-frequency two-stage octave residual block (MToRB) for feature extraction, a cascaded triple-scale feature fusion residual block (CTSFRB) for multi-scale feature integration and a multi-frequency context interaction module (MFCIM) to reduce inter-frequency correlations. Additionally, we introduce an adaptive quantization module that learns scaled uniform noise for each frequency component, enabling flexible control over quantization granularity. Furthermore, we construct a large SC image compression dataset (SDU-SCICD10K), which includes over 10,000 images spanning basic SC images, computer-rendered images, and mixed NS and SC images from both PC and mobile platforms. Experimental results demonstrate that our approach significantly improves SC image compression performance, outperforming traditional standards and state-of-the-art learning-based methods in terms of peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MS-SSIM).

图像压缩屏幕内容多频分解学习型编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。