arXiv:2605.19397eess.IVcs.MM2026-05

用感知模型优化视频传输,大幅省带宽还实时流畅。

Perception-Aware Video Semantic Communication

论文配图:Perception-Aware Video Semantic Communication
图 1 · 摘自论文原文
  • 不传运动矢量,用时空特征编码生成紧凑符号流
  • 比传统方案省75%~87%带宽,画质保持相近
  • 单模型支持自适应带宽,能在一张4090上实时运行

超高清视频流与沉浸式服务正推动无线视频流量激增。然而,在带宽受限、时延苛刻的无线链路中,传统分离式信源信道系统因侧重比特级可靠性,常在短块长传输下性能下降。此外,像素级失真优化未必符合人眼感知,而现有学习型视频编码器复杂度高且部署困难。本文提出PVSC,一种面向实时无线视频传输的感知感知视频语义通信框架。该框架摒弃显式运动矢量传输,采用时空特征编码生成紧凑且抗干扰的符号流,并设计侧信息格式化、参考缓冲管理与轻量化码率控制,实现接收端稳定重建与单一模型下的带宽自适应推理。大量实验表明,PVSC在多种数据集、分辨率、GOP配置及信道条件下均表现优异。相比工程化基准(VTM + 5G LDPC),PVSC在相近LPIPS与DISTS指标下,分别节省约75%和87%带宽,同时可在单张NVIDIA RTX 4090 GPU上实现实时推理。

原文摘要 · Abstract (English)

Ultra-high-resolution streaming and emerging immersive services are driving rapidly increasing wireless video traffic. However, perceptually pleasing video transmission over bandwidth-limited and latency-constrained wireless links remains challenging for conventional separated source-channel systems, which primarily target bit-level reliability and often suffer performance degradation under short-blocklength transmission. In addition, pixel-level distortion optimization does not necessarily align with human perception, while existing learned video codecs may incur high complexity and raise deployment issues. This paper proposes PVSC, a perception-aware video semantic communication framework for real-time wireless video transmission. PVSC eliminates explicit motion-vector transmission and exploits spatio-temporal feature coding to generate compact and channel-robust symbol streams. It also specifies side-information formatting, reference-buffer management, and lightweight rate control, enabling stable receiver-side reconstruction and bandwidth-adaptive inference with a single model. Extensive experiments demonstrate that PVSC achieves superior performance across diverse datasets, resolutions, GOP configurations, and channel conditions. Compared with the engineered ``VTM + 5G LDPC'' baseline, PVSC saves up to about 75% and 87% bandwidth at comparable LPIPS and DISTS, respectively, while enabling real-time inference on a single NVIDIA RTX 4090 GPU.

视频通信感知优化语义编码实时传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。