用大模型融合多模态语义,低带宽下也能重建高质量视频。
Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model
- 提取视频描述和首帧等语义信息,通过多模态融合重建
- 在0.0057带宽比下CLIP分数超0.92,抗噪能力强
- 适合6G沉浸式通信,尤其适用于弱信号场景
尽管传统基于香农理论的语法通信取得显著进展,但在6G沉浸式通信场景下,尤其在恶劣传输条件下仍难以满足需求。随着生成式人工智能的发展,利用高层语义信息重建视频已取得突破。本文提出一种可扩展的生成式视频语义通信框架,通过提取并传输语义信息实现高质量视频重建。发送端从源视频中提取描述文本和首帧等条件信号,分别作为文本与结构语义;接收端使用基于扩散的生成式大模型融合多模态语义进行视频重建。仿真结果表明,在超低信道带宽比(CBR=0.0057)下,该方案能有效捕捉语义信息,在不同信噪比条件下实现符合人类感知的视频重建。特别地,'首帧+描述'方案在SNR > 0 dB时,始终保持CLIP分数超过0.92,展现出强鲁棒性。
原文摘要 · Abstract (English)
Despite significant advancements in traditional syntactic communications based on Shannon's theory, these methods struggle to meet the requirements of 6G immersive communications, especially under challenging transmission conditions. With the development of generative artificial intelligence (GenAI), progress has been made in reconstructing videos using high-level semantic information. In this paper, we propose a scalable generative video semantic communication framework that extracts and transmits semantic information to achieve high-quality video reconstruction. Specifically, at the transmitter, description and other condition signals (e.g., first frame, sketches, etc.) are extracted from the source video, functioning as text and structural semantics, respectively. At the receiver, the diffusion-based GenAI large models are utilized to fuse the semantics of the multiple modalities for reconstructing the video. Simulation results demonstrate that, at an ultra-low channel bandwidth ratio (CBR), our scheme effectively captures semantic information to reconstruct videos aligned with human perception under different signal-to-noise ratios. Notably, the proposed ``First Frame+Desc." scheme consistently achieves CLIP score exceeding 0.92 at CBR = 0.0057 for SNR > 0 dB. This demonstrates its robust performance even under low SNR conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。