用大模型融合多模态数据,提升近场大阵列波束预测精度。
Structure-Aware Multimodal LLM Framework for Trustworthy Near-Field Beam Prediction
- 融合GPS、图像、激光雷达与任务提示,驱动大模型理解环境。
- 在复杂三维低空环境中实现波束预测准确率显著提升。
- 适合通信系统设计与智能感知研究者参考。
在近场超大规模多输入多输出(XL-MIMO)系统中,球面波传播将传统波束码本扩展至联合角-距离域,使常规波束训练效率极低,尤其在复杂的三维低空环境中更为严重。由于近场波束变化不仅与用户位置密切相关,还受物理环境深度影响,精准波束对准需具备深入的环境理解能力。为此,我们提出一种基于大语言模型(LLM)的多模态框架,融合历史GPS数据、RGB图像、LiDAR数据及特定任务的文本提示。通过利用LLM强大的涌现推理与泛化能力,该方法学习复杂的空间动态,实现更优的环境感知。实验表明,在真实城市低空场景下,所提方法在60%以上测试条件下将波束预测误差降低至1.8°以内,相较基线方法提升超过40%。
原文摘要 · Abstract (English)
In near-field extremely large-scale multiple-input multiple-output (XL-MIMO) systems, spherical wavefront propagation expands the traditional beam codebook into the joint angular-distance domain, rendering conventional beam training prohibitively inefficient, especially in complex 3-dimensional (3D) low-altitude environments. Furthermore, since near-field beam variations are deeply coupled not only with user positions but also with the physical surroundings, precise beam alignment demands profound environmental understanding capabilities. To address this, we propose a large language model (LLM)-driven multimodal framework that fuses historical GPS data, RGB image, LiDAR data, and strategically designed task-specific textual prompts. By utilizing the powerful emergent reasoning and generalization capabilities of the LLM, our approach learns complex spatial dynamics to achieve superior environmental comprehension...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。