arXiv:2507.06717eess.IVcs.MM2025-07被引 3

通过语义自修正机制,实现多无人机视频传输的高效低延迟与抗丢包。

QoE Optimization for Semantic Self-Correcting Video Transmission in Multi-UAV Networks

  • 将视频转为语义索引,按带宽动态发送,实现超细粒度码率调控。
  • 接收端用时空视觉变换器重建丢失信息,丢包率15%下仍保持高画质。
  • 结合强化学习联合优化资源分配,提升用户体验满意度。

实时无人机视频流对远程监控、应急响应和环境监测等时效性应用至关重要,但受限于带宽不足、延迟波动和高丢包率。为此,我们提出一种新型语义自修正视频传输框架(SSCV-G),将视频帧编码为紧凑的语义码本空间,发射端根据可用带宽自适应发送部分语义索引,实现超细粒度码率控制,提升带宽效率。接收端采用时空视觉变换器(ST-ViT)进行多帧联合解码,通过建模帧内与帧间依赖关系,恢复丢失的语义索引。为进一步提升动态网络条件下的性能,集成多用户近端策略优化(MUPPO)强化学习方案,联合优化通信资源分配与语义码率选择,以最大化用户质量体验(QoE)。大量实验表明,所提SSCV-G在编码效率、带宽适应性和丢包鲁棒性方面显著优于现有先进视频编码器;基于MUPPO的QoE优化持续超越现有基准。

原文摘要 · Abstract (English)

Real-time unmanned aerial vehicle (UAV) video streaming is essential for time-sensitive applications, including remote surveillance, emergency response, and environmental monitoring. However, it faces challenges such as limited bandwidth, latency fluctuations, and high packet loss. To address these issues, we propose a novel semantic self-correcting video transmission framework with ultra-fine bitrate granularity (SSCV-G). In SSCV-G, video frames are encoded into a compact semantic codebook space, and the transmitter adaptively sends a subset of semantic indices based on bandwidth availability, enabling fine-grained bitrate control for improved bandwidth efficiency. At the receiver, a spatio-temporal vision transformer (ST-ViT) performs multi-frame joint decoding to reconstruct dropped semantic indices by modeling intra- and inter-frame dependencies. To further improve performance under dynamic network conditions, we integrate a multi-user proximal policy optimization (MUPPO) reinforcement learning scheme that jointly optimizes communication resource allocation and semantic bitrate selection to maximize user Quality of Experience (QoE). Extensive experiments demonstrate that the proposed SSCV-G significantly outperforms state-of-the-art video codecs in coding efficiency, bandwidth adaptability, and packet loss robustness. Moreover, the proposed MUPPO-based QoE optimization consistently surpasses existing benchmarks.

视频传输无人机语义编码强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。