arXiv:2603.00482cs.CVcs.IT2026-03被引 2

提升视觉语言模型的细粒度通信能力,让图文信息更精准高效传递。

TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications

  • 双分辨率视觉分词器捕捉细节与全局特征,实现多尺度融合。
  • 引入双向注意力网络,生成紧凑视觉令牌,减少冗余信息。
  • 基于KAN的模态投影器实现非线性对齐,降低图文映射损失。

视觉-语言模型(VLM)在图像与文本理解方面具备强大能力,为智能通信提供了坚实基础。然而,其效果受限于令牌粒度不足、视觉令牌序列过长以及跨模态对齐不够等问题。为此,我们提出TaiChi框架,专为令牌通信设计。该框架采用双视觉分词器架构,分别处理高分辨率与低分辨率图像,协同捕获像素级细节与全局概念特征。引入双向注意力网络(BAN),智能融合多尺度视觉令牌,增强视觉理解并生成紧凑的视觉令牌。此外,采用基于柯尔莫哥洛夫-阿诺德网络(KAN)的模态投影器,使用可学习激活函数,实现从视觉特征到文本语义空间的精确非线性对齐,从而最小化信息损失。最后,将TaiChi集成至一个多模态多任务令牌通信系统,并采用联合VLM-通道编码方案。实验验证了TaiChi的优异性能,以及所提出的令牌通信系统的可行性与有效性。

原文摘要 · Abstract (English)

Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granularity, overlong visual token sequences, and inadequate cross-modal alignment. To overcome these challenges, we propose TaiChi, a novel VLM framework designed for token communications. TaiChi adopts a dual-visual tokenizer architecture that processes both high- and low-resolution images to collaboratively capture pixel-level details and global conceptual features. A Bilateral Attention Network (BAN) is introduced to intelligently fuse multi-scale visual tokens, thereby enhancing visual understanding and producing compact visual tokens. In addition, a Kolmogorov Arnold Network (KAN)-based modality projector with learnable activation functions is employed to achieve precise nonlinear alignment from visual features to the text semantic space, thus minimizing information loss. Finally, TaiChi is integrated into a multimodal and multitask token communication system equipped with a joint VLM-channel coding scheme. Experimental results validate the superior performance of TaiChi, as well as the feasibility and effectiveness of the TaiChi-driven token communication system.

视觉语言模型多模态通信跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。