arXiv:2511.12305cs.LG2025-11被引 4

用视觉大模型统一处理多模态无线感知,提升跨任务泛化能力。

MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing

  • 将图像、雷达、激光雷达等数据转为视觉兼容表示,实现跨模态对齐。
  • 在真实场景数据上,性能优于专用模型和通用大模型基线。
  • 适合需要统一感知框架的智能通信系统研发人员。

大型AI模型已在无线通信中用于信道建模、波束成形和资源优化。然而,现有工作大多局限于单模态输入和特定信道目标,忽视了大基础模型在统一无线感知中的潜力。为此,我们提出MMSense,一种多模态、多任务基础模型,同时解决信道中心、环境感知和人本感知任务。该框架通过将图像、雷达、激光雷达和文本数据转换为视觉兼容表示,在统一特征空间内实现有效跨模态对齐。采用模态门控机制自适应融合这些表示,以视觉为基础的大语言模型主干网络支持统一特征对齐和指令驱动的任务适配。此外,任务特定的序列注意力和基于不确定性的损失加权机制增强了跨任务泛化能力。在真实无线场景数据集上的实验表明,该方法优于任务专用和大模型基线,证实其在异构感知任务中具有强泛化能力。

原文摘要 · Abstract (English)

Large AI models have been widely adopted in wireless communications for channel modeling, beamforming, and resource optimization. However, most existing efforts remain limited to single-modality inputs and channel-specific objec- tives, overlooking the broader potential of large foundation models for unified wireless sensing. To bridge this gap, we propose MMSense, a multi-modal, multi-task foundation model that jointly addresses channel-centric, environment-aware, and human-centered sensing. Our framework integrates image, radar, LiDAR, and textual data by transforming them into vision- compatible representations, enabling effective cross-modal align- ment within a unified feature space. A modality gating mecha- nism adaptively fuses these representations, while a vision-based large language model backbone enables unified feature align- ment and instruction-driven task adaptation. Furthermore, task- specific sequential attention and uncertainty-based loss weighting mechanisms enhance cross-task generalization. Experiments on real wireless scenario datasets show that our approach outper- forms both task-specific and large-model baselines, confirming its strong generalization across heterogeneous sensing tasks.

多模态无线感知大模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。