无线传输中融合多模态语义令牌,提升车联网通信效率
AirTF: Over-the-Air Token Fusion for Task-Oriented Multi-Modal Token Communications

- 用视觉变换器提取全局语义令牌,跨传感器同步传输
- 利用信道叠加特性实现空中直接融合,频谱效率显著提升
- 基于预训练模型缓解数据不足问题,适合车载多模态场景
在车联网(IoV)中,将高维多模态感知数据传至边缘服务器以完成时效性任务时面临严重的频谱瓶颈。为此,我们提出一种面向任务的多模态令牌通信的过空气令牌融合(AirTF)框架,该框架受基础模型驱动。不同于依赖卷积神经网络(CNN)且局部感受野受限的现有分割方案,AirTF采用视觉变换器(ViT)编码器从分布式的异构传感器中提取全局上下文语义令牌。通过共享无线信道并发传输这些空间对齐的令牌,框架利用多址信道的叠加特性,直接在空中融合互补的多模态语义(如RGB与红外)。该机制相比正交传输显著提升了频谱效率。此外,预训练基础模型的引入提供了关键视觉先验,有效缓解了ViT在有限、场景特定语义分割数据集上的数据饥渴问题。实验表明,AirTF在加性高斯白噪声(AWGN)和衰落信道下均持续优于正交传输及基于CNN的融合基线。在三用户设置、残余同步误差和不完美信道状态信息估计下的额外评估进一步验证了其鲁棒性。源代码将在接收后公开。
原文摘要 · Abstract (English)
In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriented multi-modal token communications. Unlike existing schemes for segmentation that rely on convolutional neural networks (CNNs) with limited local receptive fields, AirTF leverages vision transformer (ViT) encoders to extract globally contextualized semantic tokens from distributed heterogeneous sensors. By concurrently transmitting these spatially aligned tokens over a shared wireless channel, our framework exploits the superposition property of the multiple access channel to inherently fuse complementary multi-modal semantics (e.g., RGB and infrared) directly over the air. This mechanism significantly enhances spectral efficiency compared to orthogonal transmission. Furthermore, the integration of a pre-trained foundation model provides critical visual priors, effectively addressing the data-hungry nature of ViTs on limited, scenario-specific semantic segmentation datasets. Experiments demonstrate that AirTF consistently outperforms orthogonal transmission and CNN-based fusion baselines across AWGN and fading channels. Additional evaluations under a three-user setting, residual synchronization errors, and imperfect channel state information estimation further confirm its robustness. The source code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。