arXiv:2509.03837cs.LGcs.IT2025-09中稿 · IEEE GLOBECOM 2025

用环境鸟瞰图增强大模型空间感知,提升车联网通信预测准确率

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

  • 将邻车传感数据融合生成鸟瞰图,注入大模型提供空间上下文
  • 在雨夜等恶劣条件下,预测准确率提升最高达32.7%
  • 适用于需要高鲁棒性通信预测的智能交通系统研发

精准预测车联网(V2I)通信链路质量对实现无缝切换、高效波束管理及低延迟通信至关重要。现代车辆日益丰富的传感器数据促使采用多模态大语言模型(MLLMs),因其任务适应性强且具备推理能力。然而,MLLMs缺乏三维空间理解。为此,提出一种轻量级、即插即用的鸟瞰图(BEV)注入模块:通过收集邻车传感数据构建环境BEV,并与本车输入融合,为大模型提供空间上下文。为支持真实场景下的多模态学习,开发了结合CARLA模拟器与基于MATLAB的射线追踪的协同仿真环境,生成RGB、LiDAR、GPS及无线信号数据。指令和真实答案从射线追踪输出中程序化提取。在三个V2I链路预测任务上进行大量实验:视距(LoS)与非视距(NLoS)分类、链路可用性、遮挡预测。结果表明,所提BEV注入框架在所有任务上均显著提升性能。相比仅依赖本车数据的基线,宏平均准确率最高提升13.9%;在雨夜等挑战性条件下,性能提升达32.7%,验证了该框架在恶劣环境下的鲁棒性。

原文摘要 · Abstract (English)

Accurate prediction of communication link quality metrics is essential for vehicle-to-infrastructure (V2I) systems, enabling smooth handovers, efficient beam management, and reliable low-latency communication. The increasing availability of sensor data from modern vehicles motivates the use of multimodal large language models (MLLMs) because of their adaptability across tasks and reasoning capabilities. However, MLLMs inherently lack three-dimensional spatial understanding. To overcome this limitation, a lightweight, plug-and-play bird's-eye view (BEV) injection connector is proposed. In this framework, a BEV of the environment is constructed by collecting sensing data from neighboring vehicles. This BEV representation is then fused with the ego vehicle's input to provide spatial context for the large language model. To support realistic multimodal learning, a co-simulation environment combining CARLA simulator and MATLAB-based ray tracing is developed to generate RGB, LiDAR, GPS, and wireless signal data across varied scenarios. Instructions and ground-truth responses are programmatically extracted from the ray-tracing outputs. Extensive experiments are conducted across three V2I link prediction tasks: line-of-sight (LoS) versus non-line-of-sight (NLoS) classification, link availability, and blockage prediction. Simulation results show that the proposed BEV injection framework consistently improved performance across all tasks. The results indicate that, compared to an ego-only baseline, the proposed approach improves the macro-average of the accuracy metrics by up to 13.9%. The results also show that this performance gain increases by up to 32.7% under challenging rainy and nighttime conditions, confirming the robustness of the framework in adverse settings.

车联网大模型空间感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。