arXiv:2506.14532cs.CL2025-06被引 17

用多传感器+大模型提升毫米波波束预测精度

M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models

  • 融合图像、雷达等多模态数据,借助大模型推理
  • 在标准与少样本场景下均显著优于传统深度学习模型
  • 适合车联网毫米波通信系统,尤其对传感器多样性敏感

本文提出一种名为M2BeamLLM的新型神经网络框架,用于毫米波(mmWave)大规模多输入多输出(mMIMO)通信系统的波束预测。该框架整合图像、雷达、激光雷达及GPS等多模态传感器数据,利用GPT-2等大语言模型的强大推理能力实现波束预测。通过传感数据编码、多模态对齐与融合,以及监督微调(SFT),M2BeamLLM在标准和少样本场景下均显著提升预测准确率与鲁棒性,明显优于传统深度学习模型。此外,其预测性能随感知模态多样性的增加而持续提升。本研究为车路协同(V2I)毫米波通信系统提供了高效智能的波束预测解决方案。

原文摘要 · Abstract (English)

This paper introduces a novel neural network framework called M2BeamLLM for beam prediction in millimeter-wave (mmWave) massive multi-input multi-output (mMIMO) communication systems. M2BeamLLM integrates multi-modal sensor data, including images, radar, LiDAR, and GPS, leveraging the powerful reasoning capabilities of large language models (LLMs) such as GPT-2 for beam prediction. By combining sensing data encoding, multimodal alignment and fusion, and supervised fine-tuning (SFT), M2BeamLLM achieves significantly higher beam prediction accuracy and robustness, demonstrably outperforming traditional deep learning (DL) models in both standard and few-shot scenarios. Furthermore, its prediction performance consistently improves with increased diversity in sensing modalities. Our study provides an efficient and intelligent beam prediction solution for vehicle-to-infrastructure (V2I) mmWave communication systems.

毫米波多模态大模型波束预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。