arXiv:2502.09980cs.CVcs.RO2025-02中稿 · ICRA被引 36

用多模态大模型实现车与车协同驾驶,提升感知与决策能力

V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

  • 通过多模态大模型融合多车感知信息,实现协同问答推理
  • 在新构建的V2V-QA数据集上显著优于传统融合方法
  • 适合研究智能网联汽车、自动驾驶协同系统的人士参考

当前自动驾驶车辆主要依赖自身传感器理解环境并规划路径,当传感器故障或被遮挡时可靠性下降。为解决此问题,已有基于车对车(V2V)通信的协同感知方法,但多集中于检测或跟踪任务,其对整体协同规划性能的影响仍不明确。受大语言模型在自动驾驶中应用进展启发,本文提出将多模态大模型融入协同自动驾驶的新范式,构建了V2V问答(V2V-QA)数据集与基准测试。提出基线方法V2V-LLM,利用大模型融合多辆联网自动驾驶汽车(CAVs)的感知信息,回答包括定位、显著目标识别和路径规划在内的多种驾驶相关问题。实验表明,V2V-LLM可作为统一架构完成多种协同任务,性能优于采用不同融合策略的基线方法。本工作开辟了提升未来自动驾驶系统安全性的新方向,代码与数据将公开以促进开源研究。

原文摘要 · Abstract (English)

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem, cooperative perception methods via vehicle-to-vehicle (V2V) communication have been proposed, but they have tended to focus on perception tasks like detection or tracking. How those approaches contribute to overall cooperative planning performance is still under-explored. Inspired by recent progress using Large Language Models (LLMs) to build autonomous driving systems, we propose a novel problem setting that integrates a Multimodal LLM into cooperative autonomous driving, with the proposed Vehicle-to-Vehicle Question-Answering (V2V-QA) dataset and benchmark. We also propose our baseline method Vehicle-to-Vehicle Multimodal Large Language Model (V2V-LLM), which uses an LLM to fuse perception information from multiple connected autonomous vehicles (CAVs) and answer various types of driving-related questions: grounding, notable object identification, and planning. Experimental results show that our proposed V2V-LLM can be a promising unified model architecture for performing various tasks in cooperative autonomous driving, and outperforms other baseline methods that use different fusion approaches. Our work also creates a new research direction that can improve the safety of future autonomous driving systems. The code and data will be released to the public to facilitate open-source research in this field. Our project website: https://eddyhkchiu.github.io/v2vllm.github.io/ .

协同驾驶多模态大模型V2V通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。