arXiv:2507.01728eess.SPcs.LG2025-07被引 15

用统一令牌通信提升多模态大模型的传输效率与生成能力

Token Communication in the Era of Large Models: An Information Bottleneck-Based Approach

  • 将令牌作为处理与传输的统一单元,基于生成信息瓶颈学习高效表征
  • 在动态信道下通信效率优于基线,支持跨模态可靠生成
  • 适合需要多模态理解与生成的下一代智能通信系统

本文提出UniToCom,一种统一的令牌通信范式,将令牌视为处理与无线传输的基本单位。为实现高效令牌表征,提出生成信息瓶颈(GenIB)原则,促进保留关键信息的同时支持多模态可靠生成,从而提升通信效率并降低计算复杂度。此外,设计σ-GenIB以解决自回归建模中的方差坍缩问题,保持表征多样性和稳定性。在接收端采用因果Transformer架构的多模态大语言模型(MLLM),在下一令牌预测框架下统一处理离散与连续令牌。仿真结果表明,在动态信道条件下,所提方法在性能上显著优于基线。通过融合令牌处理与MLLM,UniToCom实现了可扩展、泛化的通信能力,有利于多模态理解与生成,为下一代智能通信提供潜在解决方案。

原文摘要 · Abstract (English)

This letter proposes UniToCom, a unified token communication paradigm that treats tokens as the fundamental units for both processing and wireless transmission. Specifically, to enable efficient token representations, we propose a generative information bottleneck (GenIB) principle, which facilitates the learning of tokens that preserve essential information while supporting reliable generation across multiple modalities. By doing this, GenIB-based tokenization is conducive to improving the communication efficiency and reducing computational complexity. Additionally, we develop $σ$-GenIB to address the challenges of variance collapse in autoregressive modeling, maintaining representational diversity and stability. Moreover, we employ a causal Transformer-based multimodal large language model (MLLM) at the receiver to unify the processing of both discrete and continuous tokens under the next-token prediction paradigm. Simulation results validate the effectiveness and superiority of the proposed UniToCom compared to baselines under dynamic channel conditions. By integrating token processing with MLLMs, UniToCom enables scalable and generalizable communication in favor of multimodal understanding and generation, providing a potential solution for next-generation intelligent communications.

令牌通信多模态大模型信息瓶颈智能通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。