用多模态大模型提升自动驾驶决策能力,让系统更懂复杂路况。
Application of Multimodal Large Language Models in Autonomous Driving
- 构建虚拟问答数据集微调多模态大模型,增强对驾驶场景的理解。
- 通过思维链机制分解决策流程,在场景理解、预测与决策上表现更优。
- 适合关注智能驾驶认知推理与多模态融合的工程师与研究者。
在技术快速发展的背景下,多项前沿方法被用于提升自动驾驶(AD)系统的安全性、效率与复杂环境适应性,但其性能仍受限。为解决这一问题,我们深入研究了多模态大语言模型(MLLM)在自动驾驶中的应用。构建了虚拟问答(VQA)数据集以微调模型,改善其在驾驶任务中的表现。进一步将自动驾驶决策过程分解为场景理解、预测与决策三阶段,并引入思维链(Chain of Thought)机制,使决策过程更加清晰与准确。实验与详细分析表明,多模态大语言模型在提升自动驾驶系统认知能力方面具有重要意义。
原文摘要 · Abstract (English)
In this era of technological advancements, several cutting-edge techniques are being implemented to enhance Autonomous Driving (AD) systems, focusing on improving safety, efficiency, and adaptability in complex driving environments. However, AD still faces some problems including performance limitations. To address this problem, we conducted an in-depth study on implementing the Multi-modal Large Language Model. We constructed a Virtual Question Answering (VQA) dataset to fine-tune the model and address problems with the poor performance of MLLM on AD. We then break down the AD decision-making process by scene understanding, prediction, and decision-making. Chain of Thought has been used to make the decision more perfectly. Our experiments and detailed analysis of Autonomous Driving give an idea of how important MLLM is for AD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。