arXiv:2502.05695cs.MMcs.AI2025-02中稿 · IEEE Wireless Comm…被引 16

用潜在扩散模型压缩视频帧,省带宽还保画质。

Semantic-Aware Adaptive Video Streaming Using Latent Diffusion Models for Wireless Networks

  • 用潜在扩散模型压缩关键帧,保留语义信息
  • 结合去噪与插帧技术,提升无线环境下的画质
  • 适合5G/后5G实时视频流场景

本文提出一种新型语义通信框架,通过在FFmpeg中集成潜在扩散模型(LDMs),实现面向无线网络的实时自适应码率视频流传输。该方法解决了传统恒定码率(CBS)和自适应码率(ABS)流媒体带来的高带宽消耗、存储效率低及用户体验下降问题。通过将关键帧(I帧)压缩至潜在空间,显著降低存储与传输开销,同时保持高视觉质量;保留双向预测帧(B帧)和前向预测帧(P帧)作为调整元数据,支持用户端高效重构视频。进一步融合先进去噪与视频帧插值(VFI)技术,缓解语义模糊性,恢复帧间时序连贯性,即使在噪声干扰严重的无线环境中仍能维持高质量输出。实验表明,该方法在带宽利用与用户体验(QoE)方面优于现有最佳方案,为5G及未来后5G网络中的可扩展实时视频流提供了新路径。

原文摘要 · Abstract (English)

This paper proposes a novel Semantic Communication (SemCom) framework for real-time adaptive-bitrate video streaming by integrating Latent Diffusion Models (LDMs) within the FFmpeg techniques. This solution addresses the challenges of high bandwidth usage, storage inefficiencies, and quality of experience (QoE) degradation associated with traditional Constant Bitrate Streaming (CBS) and Adaptive Bitrate Streaming (ABS). The proposed approach leverages LDMs to compress I-frames into a latent space, offering significant storage and semantic transmission savings without sacrificing high visual quality. While retaining B-frames and P-frames as adjustment metadata to support efficient refinement of video reconstruction at the user side, the proposed framework further incorporates state-of-the-art denoising and Video Frame Interpolation (VFI) techniques. These techniques mitigate semantic ambiguity and restore temporal coherence between frames, even in noisy wireless communication environments. Experimental results demonstrate the proposed method achieves high-quality video streaming with optimized bandwidth usage, outperforming state-of-the-art solutions in terms of QoE and resource efficiency. This work opens new possibilities for scalable real-time video streaming in 5G and future post-5G networks.

视频流扩散模型语义通信5G

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。