用思维图提升车辆协同感知与决策,解决遮挡难题。
V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts
- 引入思维图机制,融合多模态大模型进行协同推理
- 在遮挡场景下,感知与预测准确率显著优于基线方法
- 适合研究智能网联汽车协同系统的学者与工程师
当前最先进的自动驾驶系统在本地传感器被道路上大型物体遮挡时可能面临安全危机。车辆间协同驾驶(V2V)被视为应对该问题的有效手段,近期框架已引入多模态大语言模型(MLLM)以整合协同感知与规划流程。然而,此前研究尚未探索将思维图(Graph-of-Thoughts)推理应用于MLLM。本文提出一种专为MLLM驱动的协同自动驾驶设计的新型思维图框架,包含提出的遮挡感知与规划感知新思路。我们构建了V2V-GoT-QA数据集,并开发了V2V-GoT模型用于训练与测试。实验表明,该方法在协同感知、预测与规划任务上均优于其他基线。项目主页:https://eddyhkchiu.github.io/v2vgot.github.io/
原文摘要 · Abstract (English)
Current state-of-the-art autonomous vehicles could face safety-critical situations when their local sensors are occluded by large nearby objects on the road. Vehicle-to-vehicle (V2V) cooperative autonomous driving has been proposed as a means of addressing this problem, and one recently introduced framework for cooperative autonomous driving has further adopted an approach that incorporates a Multimodal Large Language Model (MLLM) to integrate cooperative perception and planning processes. However, despite the potential benefit of applying graph-of-thoughts reasoning to the MLLM, this idea has not been considered by previous cooperative autonomous driving research. In this paper, we propose a novel graph-of-thoughts framework specifically designed for MLLM-based cooperative autonomous driving. Our graph-of-thoughts includes our proposed novel ideas of occlusion-aware perception and planning-aware prediction. We curate the V2V-GoT-QA dataset and develop the V2V-GoT model for training and testing the cooperative driving graph-of-thoughts. Our experimental results show that our method outperforms other baselines in cooperative perception, prediction, and planning tasks. Our project website: https://eddyhkchiu.github.io/v2vgot.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。