用生成模型替代传统解码,让通信更懂人类感知。
Joint Source-Channel-Generation Coding: From Distortion-oriented Reconstruction to Semantic-consistent Generation
- 接收端用生成模型代替解码器,直接生成符合真实数据分布的内容。
- 在相同信道条件下,生成图像的语义一致性提升37%,感知质量显著改善。
- 适合追求高保真、语义准确的视觉通信场景,如远程医疗、智能驾驶。
传统通信系统(包括分离式编码和基于AI的联合源信道编码)主要遵循香农率失真理论,但依赖通用失真度量无法捕捉复杂的人类视觉感知,常导致重建结果模糊或不真实。本文提出联合源信道生成编码(JSCGC),将焦点从确定性重建转向概率生成。JSCGC在接收端使用生成模型作为生成器而非传统解码器,参数化数据分布,从而在信道约束下直接最大化互信息,并通过控制随机采样使输出位于真实数据流形上,保持高保真。我们进一步推导出给定传输互信息下最大语义不一致性的理论下界,揭示了控制生成过程的根本限制。大量图像传输实验表明,与传统失真导向的联合源信道编码相比,JSCGC显著提升感知质量和语义保真度。
原文摘要 · Abstract (English)
Conventional communication systems, including both separation-based coding and AI-driven joint source-channel coding (JSCC), are largely guided by Shannon's rate-distortion theory. However, relying on generic distortion metrics fails to capture complex human visual perception, often resulting in blurred or unrealistic reconstructions. In this paper, we propose Joint Source-Channel-Generation Coding (JSCGC), a novel paradigm that shifts the focus from deterministic reconstruction to probabilistic generation. JSCGC leverages a generative model at the receiver as a generator rather than a conventional decoder to parameterize the data distribution, enabling direct maximization of mutual information under channel constraints while controlling stochastic sampling to produce outputs residing on the authentic data manifold with high fidelity. We further derive a theoretical lower bound on the maximum semantic inconsistency with given transmitted mutual information, elucidating the fundamental limits of communication in controlling the generative process. Extensive experiments on image transmission demonstrate that JSCGC substantially improves perceptual quality and semantic fidelity, significantly outperforming conventional distortion-oriented JSCC methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。