首个可处理多模态无线信号的通用模型,提升任务适应性。
Multimodal Wireless Foundation Models
- 融合原始基带流与图像类无线数据,统一建模不同信号形式。
- 在5项任务中表现媲美单模态模型,部分任务超越现有水平。
- 适合6G智能感知、通信与定位融合场景研究者参考。
无线基础模型(WFMs)近年来展现出强大能力,能联合执行多种无线任务并有效适应新环境。然而,现有WFMs仅处理单一模态,而实际任务中最具信息量的模态会随场景变化,无单一模态适用于所有任务。因此,应设计支持多模态输入的WFMs,以拓展任务与应用场景范围。本文提出并构建了首个可同时处理原始基带流(IQ streams)和图像类无线模态(如频谱图、信道状态信息CSI)的多模态无线基础模型,并引入面向多模态场景的掩码无线建模方法,一种自监督目标与预训练方案,用于学习来自两种模态的联合表示。我们在五项任务上进行评估:基于图像的任务(人体活动识别、射频信号分类、5G NR定位)和基于基带流的任务(射频设备指纹识别、干扰检测/分类)。结果表明,该多模态WFM性能与单模态模型相当,且在多个任务中实现超越。这证明了发展跨模态无线基础模型的巨大潜力,有助于推动面向人工智能的6G及感知-通信-定位一体化愿景。
原文摘要 · Abstract (English)
Wireless foundation models (WFMs) have recently demonstrated promising capabilities, jointly performing multiple wireless functions and adapting effectively to new environments. However, while current WFMs process only one modality, depending on the task and operating conditions, the most informative modality changes and no single modality is best for all tasks. WFMs should therefore be designed to accept multiple modalities to enable a broader and more diverse range of tasks and scenarios. In this work, we propose and build the first multimodal wireless foundation model capable of processing both raw IQ streams and image-like wireless modalities (e.g., spectrograms and CSI) and performing multiple tasks across both. We introduce masked wireless modeling for the multimodal setting, a self-supervised objective and pretraining recipe that learns a joint representation from IQ streams and image-like wireless modalities. We evaluate the model on five tasks across both modality families: image-based (human activity sensing, RF signal classification, 5G NR positioning) and IQ-based (RF device fingerprinting, interference detection/classification). The multimodal WFM is competitive with single-modality WFMs, and in several cases surpasses their performance. Our results demonstrates the strong potential of developing multimodal WFMs that support diverse wireless tasks across different modalities. We believe this provides a concrete step toward both AI-native 6G and the vision of joint sensing, communication, and localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。