解决超低码率下生成图像语义不一致问题,提升6G通信可靠性。
SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates
- 用文本、压缩图和语义残差三路引导扩散重建
- 在0.05 bpp以下使语义一致性提升23.4%(CLIC2020)
- 适合6G语义通信与低带宽图像传输场景
生成式图像压缩在近期展现出优异的感知质量,但在超低码率(bpp < 0.05)下常出现语义不一致问题,限制其在6G语义通信等带宽受限场景中的可靠应用。该问题源于引导信息不完整,导致生成过程存在语义歧义,可能产生外观自然但与源图不符的内容。本文提出语义消歧引导的生成图像压缩框架(SDGIC),在超低码率下约束基于扩散模型的重建。具体地,将源图像压缩为三个紧凑且互补的引导流:简洁文本描述全局语义、高度压缩图像(HCI)提供密集视觉证据、重建感知语义残差令牌(RSRTs)则捕捉文本与HCI条件下仍存歧义的重建相关残差语义。RSRTs直接优化于下游去噪目标,可提供源特定的语义约束以消除歧义。为有效注入三路引导,设计双路径条件扩散解码器(DPCD),通过交叉注意力实现语义条件,利用ControlNet残差提供密集视觉引导。大量实验表明,SDGIC在保持良好感知质量的同时显著提升超低码率下的语义一致性,在CLIC2020数据集上使AFINE降低23.4%。
原文摘要 · Abstract (English)
Generative image compression has recently shown impressive perceptual quality, but often suffers from semantic inconsistency at ultra-low bitrates (bpp < 0.05), limiting its reliable deployment in bandwidth-constrained scenarios such as 6G semantic communications. This inconsistency stems from incomplete guidance information, which introduces semantic ambiguity into the generation process and may lead to natural-looking but source-inconsistent content. In this work, we propose a Semantic-Disambiguation-Guided Generative Image Compression (SDGIC) framework to constrain diffusion-based reconstruction at ultra-low bitrates. Specifically, SDGIC compresses the source image into three compact and complementary guidance streams: a concise text caption for global semantics, a highly compressed image (HCI) for dense visual evidence, and Reconstruction-Aware Semantic Residual Tokens (RSRTs) for reconstruction-relevant residual semantics that remain ambiguous under the text caption and HCI conditions. The RSRTs are directly optimized toward the downstream denoising objective, enabling them to provide source-specific semantic constraints for disambiguating diffusion-based reconstruction. To inject these three guidance streams into the generation process effectively, we design a Dual-Path Conditioned Diffusion Decoder (DPCD), which uses cross-attention for semantic conditions and ControlNet residuals for dense visual guidance. Extensive experiments demonstrate that SDGIC improves semantic consistency at ultra-low bitrates while maintaining favorable perceptual quality, with a 23.4% reduction in AFINE on the CLIC2020 dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。