arXiv:2603.02470cs.ITcs.LG2026-03被引 12

用文字意图指导视频压缩,实现高效低带宽传输。

Video TokenCom: Textual Intent-Guided Multi-Rate Video Token Communications with UEP-Based Adaptive Source-Channel Coding

  • 根据用户文字意图选择性保留关键视频片段,非重点区域降精度压缩
  • 在不同信噪比下均优于传统与语义通信方法,感知与语义质量双提升
  • 适合低带宽场景下的智能视频通信,如移动实时传输

Token Communication(TokenCom)是一种新范式,受大型人工智能模型(LAMs)和多模态大语言模型(MLLMs)成功启发,将令牌作为统一的通信与计算单元,支持未来无线网络中面向语义与目标的信息高效交换。本文提出一种新型视频TokenCom框架,实现基于文本意图引导的多速率视频通信,采用不等错误保护(UEP)的源信道编码自适应。该框架融合用户意图文本、离散视频标记化与不等错误保护,以在有限带宽下提升语义保真度。首先,通过预训练视频标记器提取离散视频令牌,结合文本条件视觉-语言建模与光流传播,跨时空识别与用户意图相关的令牌。其次,提出语义感知多速率比特分配策略:高相关性意图令牌使用完整码本精度编码,非意图令牌则以降低码本精度的差分编码表示,实现速率节省并保持语义质量。最后,设计源与信道编码自适应方案,动态调整比特分配与信道编码以适应资源与链路变化。在多个视频数据集上的实验表明,所提框架在广泛信噪比范围内均优于传统及语义通信基线,在感知与语义质量上表现更优。

原文摘要 · Abstract (English)

Token Communication (TokenCom) is a new paradigm, motivated by the recent success of Large AI Models (LAMs) and Multimodal Large Language Models (MLLMs), where tokens serve as unified units of communication and computation, enabling efficient semantic- and goal-oriented information exchange in future wireless networks. In this paper, we propose a novel Video TokenCom framework for textual intent-guided multi-rate video communication with Unequal Error Protection (UEP)-based source-channel coding adaptation. The proposed framework integrates user-intended textual descriptions with discrete video tokenization and unequal error protection to enhance semantic fidelity under restrictive bandwidth constraints. First, discrete video tokens are extracted through a pretrained video tokenizer, while text-conditioned vision-language modeling and optical-flow propagation are jointly used to identify tokens that correspond to user-intended semantics across space and time. Next, we introduce a semantic-aware multi-rate bit-allocation strategy, in which tokens highly related to the user intent are encoded using full codebook precision, whereas non-intended tokens are represented through reduced codebook precision differential encoding, enabling rate savings while preserving semantic quality. Finally, a source and channel coding adaptation scheme is developed to adapt bit allocation and channel coding to varying resources and link conditions. Experiments on various video datasets demonstrate that the proposed framework outperforms both conventional and semantic communication baselines, in perceptual and semantic quality on a wide SNR range.

视频通信语义通信多速率编码意图引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。