用户指定内容可保真压缩,超低码率下仍保持高还原度
Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate
- 通过指代引导编码实现用户需求定制化压缩
- 在极低码率下仍能精准还原指定物体细节
- 适合需要可控内容保真的图像压缩场景
借助强大的生成模型,语义图像压缩(SIC)已在超低码率下取得显著进展。然而,由于视觉-语义对齐粗略且存在固有随机性,SIC在重建完全不同实例时可靠性严重不足,即使语义与原图一致。为此,我们提出一种新型指代语义图像压缩(RSIC)框架,提升用户指定内容的保真度,同时保持极端压缩比。RSIC包含三个模块:全局描述编码(GDE)、指代引导编码(RGE)和引导生成解码(GGD)。GDE与RGE分别编码全局语义信息与局部特征,GGD基于编码信息处理非均匀引导的生成过程。该方法可按用户需求灵活定制压缩,更好平衡局部保真度、全局真实感、语义对齐与比特开销。在三个数据集上的大量实验验证了所提方法的压缩效率与灵活性。
原文摘要 · Abstract (English)
With the help of powerful generative models, Semantic Image Compression (SIC) has achieved impressive performance at ultra-low bitrate. However, due to coarse-grained visual-semantic alignment and inherent randomness, the reliability of SIC is seriously concerned for reconstructing completely different object instances, even they are semantically consistent with original images. To tackle this issue, we propose a novel Referring Semantic Image Compression (RSIC) framework to improve the fidelity of user-specified content while retaining extreme compression ratios. Specifically, RSIC consists of three modules: Global Description Encoding (GDE), Referring Guidance Encoding (RGE), and Guided Generative Decoding (GGD). GDE and RGE encode global semantic information and local features, respectively, while GGD handles the non-uniformly guided generative process based on the encoded information. In this way, our RSIC achieves flexible customized compression according to user demands, which better balance the local fidelity, global realism, semantic alignment, and bit overhead. Extensive experiments on three datasets verify the compression efficiency and flexibility of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。