用语义引导的压缩技术,在极低码率下实现高质量图像重建与任务可用性。
SQ-GAN: Semantic Image Communications Using Masked Vector Quantization
- 基于语义分割生成掩码,仅编码对任务重要的图像特征。
- 在极低码率下仍保持高感知质量与语义分割准确率。
- 兼容现有系统,适合语义任务导向的图像通信场景。
本文提出语义掩码向量量化生成对抗网络(SQ-GAN),将语义驱动的图像编码与向量量化结合,优化面向任务的图像压缩。该方法仅作用于源编码,完全兼容现有系统。通过现成软件提取图像语义分割图,设计专用的语义条件自适应掩码模块(SAMM),选择性编码图像中对任务相关的特征。不同语义类别的相关性由任务决定,在训练时通过损失函数中引入相应权重实现。SQ-GAN在多个指标上超越JPEG2000、BPG及基于深度学习的先进压缩方案,尤其在极低压缩率下表现出色,显著提升重构图像的感知质量与语义分割精度。
原文摘要 · Abstract (English)
This work introduces Semantically Masked Vector Quantized Generative Adversarial Network (SQ-GAN), a novel approach integrating semantically driven image coding and vector quantization to optimize image compression for semantic/task-oriented communications. The method only acts on source coding and is fully compliant with legacy systems. The semantics is extracted from the image computing its semantic segmentation map using off-the-shelf software. A new specifically developed semantic-conditioned adaptive mask module (SAMM) selectively encodes semantically relevant features of the image. The relevance of the different semantic classes is task-specific, and it is incorporated in the training phase by introducing appropriate weights in the loss function. SQ-GAN outperforms state-of-the-art image compression schemes such as JPEG2000, BPG, and deep-learning based methods across multiple metrics, including perceptual quality and semantic segmentation accuracy on the reconstructed image, at extremely low compression rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。