arXiv:2511.21707cs.NIcs.AI2025-11被引 1

用无线信号训练多模态大模型,让网络能感知世界

Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks

  • 以无线信号为锚点,构建GPT风格的多模态模型
  • 在真实大规模数据集上表现超越现有小模型与通用大模型
  • 适合研究网络智能、感知融合的学者与工程师

大型语言模型(如ChatGPT)已在多个领域产生深远影响,具备推动网络智能化演进的巨大潜力。无线原生多模态大模型(WMLM)可通过多模态数据感知和理解物理世界,成为融合通信、感知与智能的关键使能技术,助力数亿用户实现各类智能服务。然而,当前对WMLM的研究仍处于初期阶段,针对无线网络的专用多模态大模型构建仍缺乏探索。本文梳理了WMLM的核心特征并总结现有方法,提出一种无线原生多模态训练范式。具体地,我们构建了一个GPT风格的WMLM模型,并在真实世界的大规模数据集上进行训练,利用无线信号作为对比学习的锚定模态。实验结果表明,该方法在性能上显著优于现有小规模模型及通用大模型,验证了无线信号作为通用模态的可行性,凸显了WMLM作为未来无线网络新范式的巨大潜力。

原文摘要 · Abstract (English)

Large models (LMs), such as ChatGPT, have made a significant impact across diverse domains and hold great potential to facilitate the evolution of network intelligence. Wireless-native multi-modal large models (WMLMs) can sense and understand the physical world through multi-modal data, serving as a key enabler that integrates communication, sensing, and intelligence, and thus they can boost various smart services to billions of users. However, research on WMLMs remains in its infancy, and the construction of domain-specific multi-modal large models for wireless networks is still underexplored. In this paper, we outlines the key characteristics of WMLMs and summarizes existing methods, on the basis of which a wireless-native multimodal training paradigm is proposed. Specifically, we constructed a GPT-style WMLM model and trained it on a real-world large-scale dataset, leveraging wireless signals as an anchor modality for contrastive learning. Our approach demonstrates outstanding performance compared with existing small-scale models and large multi-modal models, validating the feasibility of using wireless signals as a universal modality and highlighting WMLM's potential to emerge as a new paradigm for future wireless networks.

多模态模型无线感知网络智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。