用扩散模型实现无损图像传输,提升噪声信道下的恢复精度。
Adapting Diffusion Language Models for Lossless Pixel-Level Image Transmission

- 基于扩散语言模型,通过双向注意力同步解码多个像素令牌。
- 在加性高斯白噪声和瑞利衰落信道下,重建误差更低,性能超越基线。
- 适合对图像无损传输有要求的通信系统设计者参考。
无损像素级图像传输是超越语义通信的基础范式,因其精确恢复需兼顾准确的符号概率建模与在噪声信道中的可靠传输。本文提出DDM-SSCC,一种基于离散扩散模型的分离源信道编码框架,用于无损图像传输。不同于按栅格顺序的自回归编码,该源编解码器将扩散语言模型用于像素令牌恢复,并在双向注意力下执行同步反向算术编码,可在一次反向去噪步骤中编码多个被掩码的令牌。此渐进式恢复过程生成更利于噪声传输的源表示,因为新恢复的令牌可作为后续去噪步骤中的双向上下文。为弥合生成导向的掩码去噪与无损算术编码之间的差距,我们引入哈尔顿引导的去噪顺序、掩码比例感知的余弦调度以及轻量级温度校准模块。这些设计分别提升了空间覆盖范围,使去噪速率适应上下文可靠性,并校准算术编码所用的概率表。在CIFAR10、DIV2K-LR-X4和Kodak数据集上,于加性白高斯噪声和瑞利衰落信道下的实验表明,DDM-SSCC在精确恢复性能上优于代表性无损与语义通信基线;消融实验验证了所提去噪顺序、调度和校准模块的有效性。
原文摘要 · Abstract (English)
Lossless pixel-level image transmission is a fundamental regime beyond semantic communications, because exact recovery requires both accurate symbol probability modeling and reliable delivery over noisy channels. This paper proposes DDM-SSCC, a discrete-diffusion-model-based separate source-channel coding framework for lossless image transmission. Different from raster-order autoregressive coding, the proposed source codec adapts a diffusion language model to pixel-token restoration and performs synchronized reverse arithmetic coding under bidirectional attention, allowing multiple masked tokens to be coded within one reverse denoising step. This progressive restoration process also yields a more favorable source representation for noisy transmission, since newly restored tokens can serve as bidirectional context in subsequent denoising steps. To bridge the gap between generation-oriented masked denoising and lossless arithmetic coding, we further introduce a Halton-guided denoising order, a mask-ratio-aware cosine schedule, and a lightweight temperature calibration module. These designs respectively improve spatial coverage, adapt the denoising pace to context reliability, and calibrate the probability tables used by arithmetic coding. Experiments on CIFAR10, DIV2K-LR-X4, and Kodak over additive white Gaussian noise and Rayleigh fading channels show that DDM-SSCC achieves better exact-recovery performance than representative lossless and semantic communication baselines, while ablation studies verify the effectiveness of the proposed denoising order, schedule, and calibration modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。