用大模型拆解语义,低比特压缩也能保清晰和准确。
Extremely low-bitrate Image Compression Semantically Disentangled by LMMs from a Human Perception Perspective
- 分步恢复:先压缩参考图,再用大模型提取语义信息逐步还原。
- 在0.05 bpp以下仍保持高感知质量与语义一致。
- 适合需要极低带宽却要保关键信息的场景,如远程医疗、卫星图像。
在极低比特率下同时实现语义一致性和高感知质量仍是图像压缩的重大挑战。受人类渐进式感知机制启发,本文提出一种语义解耦图像压缩框架(SEDIC)。首先通过学习的图像编码器获得极度压缩的参考图像;随后利用大模型(LMMs)提取整体描述、物体细节描述及语义分割掩码等核心语义成分。提出无需训练的注意力引导物体修复模型(ORAG),基于预训练ControlNet,结合物体级文本描述和语义掩码来恢复物体细节。在此基础上设计多阶段语义解码器,从极度压缩的参考图像出发,逐个物体渐进式恢复细节,最终生成高质量、高保真的重建结果。实验表明,SEDIC在极低比特率(≤0.05 bpp)下显著优于现有方法,显著提升感知质量和语义一致性。
原文摘要 · Abstract (English)
It remains a significant challenge to compress images at extremely low bitrate while achieving both semantic consistency and high perceptual quality. Inspired by human progressive perception mechanism, we propose a Semantically Disentangled Image Compression framework (SEDIC) in this paper. Initially, an extremely compressed reference image is obtained through a learned image encoder. Then we leverage LMMs to extract essential semantic components, including overall descriptions, object detailed description, and semantic segmentation masks. We propose a training-free Object Restoration model with Attention Guidance (ORAG) built on pre-trained ControlNet to restore object details conditioned by object-level text descriptions and semantic masks. Based on the proposed ORAG, we design a multistage semantic image decoder to progressively restore the details object by object, starting from the extremely compressed reference image, ultimately generating high-quality and high-fidelity reconstructions. Experimental results demonstrate that SEDIC significantly outperforms state-of-the-art approaches, achieving superior perceptual quality and semantic consistency at extremely low-bitrates ($\le$ 0.05 bpp).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。