arXiv:2502.20927eess.IV2025-02被引 22

用生成式AI实现低延迟高清视频无线传输,抗噪能力强。

Goal-Oriented Semantic Communication for Wireless Video Transmission via Generative AI

  • 基于扩散模型设计语义编码解码框架,压缩视频数据并重建高质量画面。
  • 在已知信道下,PSNR提升最高达69%,误差降低超50%。
  • 适用于带宽受限、噪声严重的无线场景,适合智能终端实时视频通信。

高效视频传输对视觉驱动的数字协作至关重要。为在带宽受限且存在噪声的无线信道中实现低延迟、高质量视频传输,本文提出一种基于稳定扩散(SD)的目标导向语义通信(GSC)框架。首先设计语义编码器,从视频中提取关键帧和相关语义信息(SI),显著减少传输数据量;随后开发语义解码器,基于接收的SI重建关键帧,并通过帧插值生成完整视频以保证画质。针对无线信道噪声影响,提出条件于瞬时信道增益的SD去噪器(SD-GSC),在已知信道下可有效去除噪声。对于未知信道,进一步提出并行SD去噪器(PSD-GSC),联合学习信道增益分布并完成去噪。实验表明,在已知信道下,SD-GSC相比ADJSCC、Latent-Diff DNSC、DeepWiVe和DVST,PSNR分别提升69%、58%、33%、38%,均方误差(MSE)降低52%、50%、41%、45%,弗雷谢特视频距离(FVD)减少38%、32%、22%、24%。在未知信道下,PSD-GSC相较MMSE均衡增强的SD-GSC,PSNR提升17%,MSE下降29%,FVD降低19%。这些显著性能提升证明了所提方法在不同信道条件下的鲁棒性与优越性。

原文摘要 · Abstract (English)

Efficient video transmission is essential for seamless communication and collaboration within the visually-driven digital landscape. To achieve low latency and high-quality video transmission over a bandwidth-constrained noisy wireless channel, we propose a stable diffusion (SD)-based goal-oriented semantic communication (GSC) framework. In this framework, we first design a semantic encoder that effectively identify the keyframes from video and extract the relevant semantic information (SI) to reduce the transmission data size. We then develop a semantic decoder to reconstruct the keyframes from the received SI and further generate the full video from the reconstructed keyframes using frame interpolation to ensure high-quality reconstruction. Recognizing the impact of wireless channel noise on SI transmission, we also propose an SD-based denoiser for GSC (SD-GSC) condition on an instantaneous channel gain to remove the channel noise from the received noisy SI under a known channel. For scenarios with an unknown channel, we further propose a parallel SD denoiser for GSC (PSD-GSC) to jointly learn the distribution of channel gains and denoise the received SI. It is shown that, with the known channel, our proposed SD-GSC outperforms state-of-the-art ADJSCC, Latent-Diff DNSC, DeepWiVe and DVST, improving Peak Signal-to-Noise Ratio (PSNR) by 69%, 58%, 33% and 38%, reducing mean squared error (MSE) by 52%, 50%, 41% and 45%, and reducing Fréchet Video Distance (FVD) by 38%, 32%, 22% and 24%, respectively. With the unknown channel, our PSD-GSC achieves a 17% improvement in PSNR, a 29% reduction in MSE, and a 19% reduction in FVD compared to MMSE equalizer-enhanced SD-GSC. These significant performance improvements demonstrate the robustness and superiority of our proposed methods in enhancing video transmission quality and efficiency under various channel conditions.

视频传输生成式AI语义通信无线信道

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。