用生成模型提升无线图像传输的清晰度与视觉质量
U-Net-Based Generative Joint Source-Channel Coding for Wireless Image Transmission
- 采用U-Net结构的解码器融合多尺度特征,提升重建质量
- 结合结构相似性和均方误差损失,优化像素级还原效果
- 对抗训练增强低分辨率图像鲁棒性,适合实际通信场景
基于深度学习的联合信源信道编码(JSCC)在无线图像传输中取得显著进展。然而,现有方法或仅关注传统失真指标而无法保证高感知质量,或计算复杂度较高。本文提出两种基于深度生成架构的深度联合信源信道编码(DeepJSCC)方法。首先提出G-UNet-JSCC,其解码器采用带跳跃连接的U-Net结构,通过融合高低层特征提升重建图像的像素级保真度与感知质量。为进一步提升像素精度,采用结构相似性(SSIM)与均方误差(MSE)加权和进行编解码端联合优化。在此基础上,提出cGAN-JSCC,通过对抗训练增强解码器:保留G-UNet-JSCC的编码器,对解码器生成器与基于块的判别器进行两阶段训练——外层端到端优化使用MSE损失,内层联合最小化对抗损失与失真损失。仿真结果表明,所提方法在高、低分辨率图像上均实现优异的像素保真度与感知质量。对于低分辨率图像,cGAN-JSCC相比G-UNet-JSCC具有更优的重建性能与更强的信道变化鲁棒性。
原文摘要 · Abstract (English)
Deep learning (DL)-based joint source-channel coding (JSCC) methods have achieved remarkable success in wireless image transmission. However, these methods either focus on conventional distortion metrics that do not necessarily yield high perceptual quality or incur high computational complexity. In this paper, we propose two DL-based JSCC (DeepJSCC) methods that leverage deep generative architectures for wireless image transmission. Specifically, we propose G-UNet-JSCC, a scheme comprising an encoder and a U-Net-based generator serving as the decoder. Its skip connections enable multi-scale feature fusion to improve both pixel-level fidelity and perceptual quality of reconstructed images by integrating low- and high-level features. To further enhance pixel-level fidelity, the encoder and the U-Net-based decoder are jointly optimized using a weighted sum of structural similarity and mean-squared error (MSE) losses. Building upon G-UNet-JSCC, we further develop a DeepJSCC method called cGAN-JSCC, where the decoder is enhanced through adversarial training. In this scheme, we retain the encoder of G-UNet-JSCC and adversarially train the decoder's generator against a patch-based discriminator. cGAN-JSCC employs a two-stage training procedure. The outer stage trains the encoder and the decoder end-to-end using an MSE loss, while the inner stage adversarially trains the decoder's generator and the discriminator by minimizing a joint loss combining adversarial and distortion losses. Simulation results demonstrate that the proposed methods achieve superior pixel-level fidelity and perceptual quality on both high- and low-resolution images. For low-resolution images, cGAN-JSCC achieves better reconstruction performance and greater robustness to channel variations than G-UNet-JSCC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。