提出多尺度时空网络,实现无线视频端到端高效传输
A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission
- 设计多尺度视觉Transformer编码解码器,捕捉长期帧时空特征
- 动态选择关键语义令牌,按需调整编码长度,节省带宽
- 相比传统分步编码和现有深度联合编码,重建质量更优
深度联合源信道编码(DeepJSCC)在文本、语音和图像的语义通信中展现出潜力。然而,无线视频传输因难以提取并紧凑表示空间与时间特征,且对带宽和计算资源要求高,面临更大挑战。为此,我们提出一种新型视频深度联合源信道编码(VDJSCC)方法,实现无线信道上的端到端视频传输。该方法设计了多尺度视觉Transformer编码器与解码器,有效捕捉长期帧的时空表征;同时提出动态令牌选择模块,屏蔽空间或时间维度中语义重要性较低的令牌,通过调节保留比例实现内容自适应的可变长度视频编码。实验结果表明,与采用分离源码和信道码的数字方案及其他DeepJSCC方法相比,本方法在重建质量与带宽降低方面均表现更优。
原文摘要 · Abstract (English)
Deep joint source-channel coding (DeepJSCC) has shown promise in wireless transmission of text, speech, and images within the realm of semantic communication. However, wireless video transmission presents greater challenges due to the difficulty of extracting and compactly representing both spatial and temporal features, as well as its significant bandwidth and computational resource requirements. In response, we propose a novel video DeepJSCC (VDJSCC) approach to enable end-to-end video transmission over a wireless channel. Our approach involves the design of a multi-scale vision Transformer encoder and decoder to effectively capture spatial-temporal representations over long-term frames. Additionally, we propose a dynamic token selection module to mask less semantically important tokens from spatial or temporal dimensions, allowing for content-adaptive variable-length video coding by adjusting the token keep ratio. Experimental results demonstrate the effectiveness of our VDJSCC approach compared to digital schemes that use separate source and channel codes, as well as other DeepJSCC schemes, in terms of reconstruction quality and bandwidth reduction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。