arXiv:2601.03112eess.IVcs.CV2026-01被引 1

用扩散模型提升极端信道下的图像传输质量与语义一致性

DiT-JSCC: Rethinking Deep JSCC with Diffusion Transformers and Semantic Representations

  • 设计双分支编码器与分层生成解码器,兼顾语义与细节
  • 在极低带宽下保持高语义一致性和视觉质量
  • 无需训练的带宽自适应策略提升传输效率

生成式联合源信道编码(GJSCC)作为新一代深度联合源信道编码范式,在超低带宽和低信噪比等极端无线信道条件下实现了高保真、鲁棒的图像传输。现有方法多采用扩散模型作为生成解码器,但常出现视觉逼真而语义不一致的问题。这源于重建导向的编码器与生成解码器之间的根本性不匹配:前者缺乏显式语义判别能力,无法提供可靠的条件线索。本文提出DiT-JSCC,一种新型GJSCC框架,可联合学习以语义优先的表示编码器与基于扩散变换器(DiT)的生成解码器。具体而言,设计了语义-细节双分支编码器,与粗到细的条件式DiT解码器自然对齐,优先保障极端信道下的语义一致性。此外,引入受科尔莫戈罗夫复杂度启发的训练无关自适应带宽分配策略,进一步提升传输效率,重新定义生成解码时代的“信息价值”概念。大量实验表明,DiT-JSCC在语义一致性和视觉质量上均持续优于现有方法,尤其在极端条件下表现突出。

原文摘要 · Abstract (English)

Generative joint source-channel coding (GJSCC) has emerged as a new Deep JSCC paradigm for achieving high-fidelity and robust image transmission under extreme wireless channel conditions, such as ultra-low bandwidth and low signal-to-noise ratio. Recent studies commonly adopt diffusion models as generative decoders, but they frequently produce visually realistic results with limited semantic consistency. This limitation stems from a fundamental mismatch between reconstruction-oriented JSCC encoders and generative decoders, as the former lack explicit semantic discriminability and fail to provide reliable conditional cues. In this paper, we propose DiT-JSCC, a novel GJSCC backbone that can jointly learn a semantics-prioritized representation encoder and a diffusion transformer (DiT) based generative decoder, our open-source project aims to promote the future research in GJSCC. Specifically, we design a semantics-detail dual-branch encoder that aligns naturally with a coarse-to-fine conditional DiT decoder, prioritizing semantic consistency under extreme channel conditions. Moreover, a training-free adaptive bandwidth allocation strategy inspired by Kolmogorov complexity is introduced to further improve the transmission efficiency, thereby indeed redefining the notion of information value in the era of generative decoding. Extensive experiments demonstrate that DiT-JSCC consistently outperforms existing JSCC methods in both semantic consistency and visual quality, particularly in extreme regimes.

生成式编码扩散模型语义一致性无线传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。