arXiv:2509.05925cs.CVcs.IT2025-09被引 7

用CLIP模型压缩图像语义,比特率低于主流方法5%。

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

  • 将图像转为CLIP特征向量压缩,专注语义保留而非像素还原。
  • 平均每像素仅需2-3×10⁻³比特,不足主流方案的5%。
  • 极端压缩下仍能跨数据分布和任务零样本保持鲁棒性。

基于深度学习的有损图像压缩方法通过端到端训练和先进架构实现了良好的率失真性能。然而,新兴应用越来越关注语义保全而非像素级重建,并要求在多样数据分布和下游任务中具备鲁棒表现。这催生了更先进的语义压缩范式。受多模态基础模型零样本能力和表征能力的启发,我们提出一种基于对比语言-图像预训练(CLIP)模型的新语义压缩方法。该方法不以重建图像为目标,而是将CLIP特征嵌入压缩至最少比特数,同时保持不同任务间的语义信息。实验表明,该方法在基准数据集上维持了良好的语义完整性,平均比特率为约2-3×10⁻³比特/像素,不足主流图像压缩方法实现相当性能所需比特率的5%。尤为显著的是,即使在极端压缩条件下,该方法仍表现出对多样化数据分布和下游任务的零样本鲁棒性。

原文摘要 · Abstract (English)

Recent deep learning-based methods for lossy image compression achieve competitive rate-distortion performance through extensive end-to-end training and advanced architectures. However, emerging applications increasingly prioritize semantic preservation over pixel-level reconstruction and demand robust performance across diverse data distributions and downstream tasks. These challenges call for advanced semantic compression paradigms. Motivated by the zero-shot and representational capabilities of multimodal foundation models, we propose a novel semantic compression method based on the contrastive language-image pretraining (CLIP) model. Rather than compressing images for reconstruction, we propose compressing the CLIP feature embeddings into minimal bits while preserving semantic information across different tasks. Experiments show that our method maintains semantic integrity across benchmark datasets, achieving an average bit rate of approximately 2-3* 10(-3) bits per pixel. This is less than 5% of the bitrate required by mainstream image compression approaches for comparable performance. Remarkably, even under extreme compression, the proposed approach exhibits zero-shot robustness across diverse data distributions and downstream tasks.

语义压缩CLIP多模态低比特

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。