用扩散模型修复压缩失真图像,提升极端条件下的视觉质量。
SING: Semantic Image Communications using Null-Space and INN-Guided Diffusion Models
- 两阶段框架将图像恢复建模为逆问题,结合空域与可逆网络建模退化过程。
- 在低带宽、低信噪比下仍保持高感知质量,显著优于传统方法。
- 适用于训练测试分布不一致场景,适合通信系统中对视觉体验要求高的应用。
基于深度神经网络的联合源信道编码(DeepJSCC)在无线图像传输中表现出色。现有方法多关注重建图像与原始图像间的失真最小化,常忽视感知质量,在极低带宽压缩比(BCR)和低信噪比(SNR)条件下易导致严重感知劣化。本文提出SING,一种新型两阶段JSCC框架,将从受损重建中恢复高质量源图像建模为逆问题。根据接收端是否掌握编码器/解码器及信道信息,SING可将随机退化近似为线性变换,或利用可逆神经网络(INN)进行精确建模。两种方式均支持扩散模型无缝融入重建流程,显著提升感知质量。实验表明,SING在极端条件下(如训练测试数据分布显著不匹配)仍优于DeepJSCC及其他方法,实现更优的视觉效果。
原文摘要 · Abstract (English)
Joint source-channel coding systems based on deep neural networks (DeepJSCC) have recently demonstrated remarkable performance in wireless image transmission. Existing methods primarily focus on minimizing distortion between the transmitted image and the reconstructed version at the receiver, often overlooking perceptual quality. This can lead to severe perceptual degradation when transmitting images under extreme conditions, such as low bandwidth compression ratios (BCRs) and low signal-to-noise ratios (SNRs). In this work, we propose SING, a novel two-stage JSCC framework that formulates the recovery of high-quality source images from corrupted reconstructions as an inverse problem. Depending on the availability of information about the DeepJSCC encoder/decoder and the channel at the receiver, SING can either approximate the stochastic degradation as a linear transformation, or leverage invertible neural networks (INNs) for precise modeling. Both approaches enable the seamless integration of diffusion models into the reconstruction process, enhancing perceptual quality. Experimental results demonstrate that SING outperforms DeepJSCC and other approaches, delivering superior perceptual quality even under extremely challenging conditions, including scenarios with significant distribution mismatches between the training and test data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。