arXiv:2608.21651eess.IV2026-08

用流匹配提升无线图像传输质量,低信噪比下仍保持清晰细节。

FlowSem: Flow Matching for Adaptive Wireless Image Transmission in Semantic Communication

论文配图:FlowSem: Flow Matching for Adaptive Wireless Image Transmission in Semantic Communication
图 1 · 摘自论文原文
  • 分两阶段:先压缩编码,再用条件流匹配生成图像
  • 低信噪比下FID比扩散模型低60%,重建更保真
  • 只需少量迭代步数,适合实时通信场景

在恶劣信道条件和严格带宽限制下,无线图像传输需同时保证像素级保真度和有意义的视觉结构。传统分离式系统(如BPG+LDPC)存在悬崖效应,而深度联合源信道编码(DeepJSCC)虽可渐进退化,但强压缩与严重干扰下会丢失细节。本文提出两阶段流匹配语义通信框架FlowSem:第一阶段为信噪比自适应的DeepJSCC模型,将图像映射为信道符号并生成粗略重构;第二阶段采用条件流匹配模型,从高斯噪声中生成最终图像,条件为DeepJSCC重构结果和信道SNR。在Cityscapes数据集上,于加性白高斯噪声(AWGN)和瑞利衰落信道下评估,固定信道符号预算。对比基线包括率匹配的BPG+LDPC、DeepJSCC、去噪扩散概率模型(DDPM)和去噪扩散隐式模型(DDIM)。结果表明,FlowSem在不同信道条件下均实现竞争性像素级保真度,并优于生成类基线的结构与感知重建质量。在低信噪比下,其弗雷谢特初始距离(FID)较扩散基线降低最高达60%。此外,仅需少数常微分方程(ODE)积分步数即可达到高质量重构,相比标准DDPM和加速版DDIM采样,具有更优的质量-延迟权衡。

原文摘要 · Abstract (English)

Wireless image transmission becomes challenging under poor channel conditions and stringent bandwidth constraints, as the receiver needs to preserve both pixel-level fidelity and meaningful visual structure. Classical separation-based systems, such as better portable graphics with low-density parity-check coding (BPG+LDPC), may suffer from cliff-effect behavior. While deep joint source-channel coding (DeepJSCC) provides graceful degradation as channel conditions worsen, its reconstructions may lose fine details under strong compression and severe channel distortion. To address this limitation, this paper proposes a two-stage flow matching-based semantic communication framework, termed FlowSem. In the first stage, a signal-to-noise ratio (SNR)-adaptive DeepJSCC model maps the source image into channel symbols and produces a coarse reconstruction at the receiver. In the second stage, a conditional flow matching model generates the final image from Gaussian noise conditioned on the DeepJSCC reconstruction and channel SNR. The proposed framework is evaluated on the Cityscapes dataset under additive white Gaussian noise (AWGN) and Rayleigh fading channels using a fixed channel-symbol budget. The baselines include a rate-matched BPG+LDPC system, DeepJSCC, a denoising diffusion probabilistic model (DDPM), and a denoising diffusion implicit model (DDIM). Results show that FlowSem achieves competitive pixel-level fidelity and improved structural and perceptual reconstruction quality over the considered generative baselines across different channel conditions. FlowSem provides up to 60% lower Fréchet Inception Distance (FID) than the diffusion baselines at low SNRs. Moreover, FlowSem reaches high reconstruction quality using only a few ODE integration steps, providing a favorable quality-latency tradeoff compared with standard DDPM and accelerated DDIM sampling.

语义通信流匹配图像传输无线通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。