arXiv:2504.01279stat.APeess.IV2025-04中稿 · ICME2025被引 2

用文字描述提升图像压缩质量,兼顾速度与效果。

SELIC: Semantic-Enhanced Learned Image Compression via High-Level Textual Guidance

  • 通过文本编码器提取图像语义,融合到压缩特征中
  • 在多种比特率下比基线模型提升0.1-0.15 dB PSNR
  • 仅增加少量计算开销,适合实际应用

学习型图像压缩(LIC)已取得显著进展,但如何有效融入高层语义信息仍是挑战。本文提出一种名为SELIC的语义增强学习图像压缩框架,利用高阶文本引导提升率失真性能。具体而言,SELIC通过文本编码器从输入图像中提取丰富的语义描述,并将其转换为固定维度张量,与图像隐含表示无缝融合。将该融合张量直接嵌入压缩流程,不需解码端额外输入,保持快速高效解码。在标准数据集(如Kodak)上的大量实验表明,引入语义信息显著提升压缩质量。所提方法在不同比特率下相较无语义集成的基线模型平均提升0.1–0.15 dB PSNR,且相比VVC实现4.9%的BD-rate改进。该提升仅带来微小计算开销,使SELIC成为先进图像压缩应用的实用解决方案。

原文摘要 · Abstract (English)

Learned image compression (LIC) techniques have achieved remarkable progress; however, effectively integrating high-level semantic information remains challenging. In this work, we present a \underline{S}emantic-\underline{E}nhanced \underline{L}earned \underline{I}mage \underline{C}ompression framework, termed \textbf{SELIC}, which leverages high-level textual guidance to improve rate-distortion performance. Specifically, \textbf{SELIC} employs a text encoder to extract rich semantic descriptions from the input image. These textual features are transformed into fixed-dimension tensors and seamlessly fused with the image-derived latent representation. By embedding the \textbf{SELIC} tensor directly into the compression pipeline, our approach enriches the bitstream without requiring additional inputs at the decoder, thereby maintaining fast and efficient decoding. Extensive experiments on benchmark datasets (e.g., Kodak) demonstrate that integrating semantic information substantially enhances compression quality. Our \textbf{SELIC}-guided method outperforms a baseline LIC model without semantic integration by approximately 0.1-0.15 dB across a wide range of bit rates in PSNR and achieves a 4.9\% BD-rate improvement over VVC. Moreover, this improvement comes with minimal computational overhead, making the proposed \textbf{SELIC} framework a practical solution for advanced image compression applications.

图像压缩语义增强学习型压缩文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。