arXiv:2601.07512cs.LGeess.IV2026-01被引 3

用流匹配实现低延迟无线图像解码,让生成式解码更快更可靠。

Land-then-transport: A Flow Matching-Based Generative Decoder for Wireless Image Transmission

  • 提出新型‘先落地再传输’范式,将无线信道融入连续概率流中
  • 仅需少量微分方程步数即达高感知质量,比扩散模型快得多
  • 适用于高斯、瑞利和MIMO信道,无需重新训练

由于速率与可靠性要求严格,传统分层设计和联合源信道编码(JSCC)在低时延场景下仍面临挑战。基于扩散的生成解码器虽能通过学习图像先验获得优异感知质量,但迭代随机去噪导致解码延迟高。为实现低延迟解码,本文提出一种在新“先落地再传输”(LTT)范式下的流匹配(FM)生成解码器,将物理无线信道紧密嵌入连续时间概率流中。针对加性高斯白噪声(AWGN)信道,构建高斯平滑路径,其噪声调度对应有效噪声水平,并推导出闭式教师速度场;通过条件流匹配训练神经网络学生速度场,得到确定性、信道感知的常微分方程(ODE)解码器,计算复杂度与ODE步数呈线性关系。推理时仅需估计有效噪声方差以设定ODE起始时间。进一步证明,瑞利衰落和多输入多输出(MIMO)信道可通过线性最小均方误差(MMSE)均衡及奇异值域处理映射为等效高斯信道,且起始时间可校准。因此,同一概率路径与训练速度场可复用于瑞利与MIMO信道而无需重训。在MNIST、Fashion-MNIST和DIV2K数据集上,于AWGN、瑞利及MIMO信道的实验表明,本方法持续优于JPEG2000+LDPC、DeepJSCC及扩散基线,且仅用少数ODE步数即可实现良好感知质量。

原文摘要 · Abstract (English)

Due to strict rate and reliability demands, wireless image transmission remains difficult for both classical layered designs and joint source-channel coding (JSCC), especially under low latency. Diffusion-based generative decoders can deliver strong perceptual quality by leveraging learned image priors, but iterative stochastic denoising leads to high decoding delay. To enable low-latency decoding, we propose a flow-matching (FM) generative decoder under a new land-then-transport (LTT) paradigm that tightly integrates the physical wireless channel into a continuous-time probability flow. For AWGN channels, we build a Gaussian smoothing path whose noise schedule indexes effective noise levels, and derive a closed-form teacher velocity field along this path. A neural-network student vector field is trained by conditional flow matching, yielding a deterministic, channel-aware ODE decoder with complexity linear in the number of ODE steps. At inference, it only needs an estimate of the effective noise variance to set the ODE starting time. We further show that Rayleigh fading and MIMO channels can be mapped, via linear MMSE equalization and singular-value-domain processing, to AWGN-equivalent channels with calibrated starting times. Therefore, the same probability path and trained velocity field can be reused for Rayleigh and MIMO without retraining. Experiments on MNIST, Fashion-MNIST, and DIV2K over AWGN, Rayleigh, and MIMO demonstrate consistent gains over JPEG2000+LDPC, DeepJSCC, and diffusion-based baselines, while achieving good perceptual quality with only a few ODE steps. Overall, LTT provides a deterministic, physically interpretable, and computation-efficient framework for generative wireless image decoding across diverse channels.

生成式解码流匹配无线传输低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。