用上下文模型提升无线传输中令牌的可靠性和效率
Context-Aware Wireless Token Communication via Joint Token Masking and Detection

- 利用掩码语言模型共享上下文,实现收发端协同理解
- 在噪声信道下重建性能提升1.77倍(Europarl)和1.63倍(WikiText-103)
- 适合通信与NLP融合场景,尤其关注资源受限下的高效传输
语言驱动应用中基于令牌的表示日益普及,推动了无线令牌通信的发展,其中令牌被视为传输的基本单元。然而,传统通信系统忽视令牌间的依赖关系,且均匀分配传输资源,在信道恶化条件下导致无线资源利用效率低下。本文提出一种上下文感知的令牌通信框架,以掩码语言模型(MLM)作为收发端共享的上下文模型。接收端设计了一种上下文感知的令牌检测方法,基于贝叶斯框架将信道似然与MLM生成的上下文先验相结合,实现对噪声信道中令牌的鲁棒推断。发送端提出一种上下文感知的令牌掩码策略,选择性跳过接收端可可靠推断的令牌,从而将有限功率集中于更关键的令牌上。上述组件通过共享的MLM实现联合设计,构建了统一的收发端高效令牌传输与检测框架。仿真结果表明,该框架相比传统及现有令牌通信方案显著提升了重建性能,在Europarl语料库和WikiText-103数据集上分别实现了最高1.77倍和1.63倍的性能增益。
原文摘要 · Abstract (English)
The increasing use of token-based representations in language-driven applications has motivated wireless token communication, where tokens are treated as fundamental units for transmission. However, conventional communication systems overlook dependencies among tokens and allocate transmission resources uniformly, leading to inefficient use of limited wireless resources under channel impairments. In this paper, we propose a context-aware token communication framework that leverages a masked language model (MLM) as a shared contextual model between the transmitter (Tx) and receiver (Rx). At the Rx, we develop a context-aware token detection method that integrates channel likelihoods with MLM-based contextual priors under a Bayesian formulation, enabling robust token inference over noisy channels. At the Tx, we propose a context-aware token masking strategy that selectively omits tokens that can be reliably inferred at the Rx, allowing the available power budget to be concentrated on more informative tokens. These components are jointly designed through a shared MLM, establishing a unified Tx-Rx framework for efficient token transmission and detection. Simulation results demonstrate that the proposed framework significantly improves reconstruction performance compared to conventional and existing token communication schemes, achieving up to 1.77X and 1.63X performance gains on the Europarl corpus and WikiText-103 datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。