arXiv:2503.10621cs.CVcs.RO2025-03被引 49

构建驾驶场景多步推理数据集与模型,提升自动驾驶理解能力。

DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

  • 设计18000+题的驾驶场景多步推理数据集,覆盖感知、预测、规划任务
  • 自研大模型在准确率上比最优开源模型高7.49%,推理得分提升3.62%
  • 适合研究自动驾驶认知推理、多模态模型评估的学者与工程师

尽管大型多模态模型(LMMs)在各类视觉问答(VQA)任务中表现强劲,但某些任务仍需复杂多步推理才能得出准确答案。自动驾驶是其中最具挑战性的领域之一,需对视觉线索进行序列化、解释性理解,以支持有效感知、预测与规划。然而,现有常见VQA基准通常只关注最终答案的准确性,忽视生成答案背后的推理过程。此外,缺乏针对真实驾驶场景下多步推理的系统性评估框架。为此,我们提出DriveLMM-o1,一个专为推进自动驾驶多步视觉推理而设计的新数据集与基准。该基准包含训练集超过18000个VQA样本,测试集超4000个,涵盖感知、预测与规划类问题,每个问题均配有逐步推理路径,确保逻辑一致性。我们进一步基于该数据集微调一个大型多模态模型,在复杂驾驶场景中表现稳健。同时,我们在新数据集上对比多种开源与闭源方法,系统评估其推理能力。所提模型相较此前最优开源模型,最终答案准确率提升7.49%,推理得分提高3.62%。相关框架、数据集与模型已公开于https://github.com/ayesha-ishaq/DriveLMM-o1。

原文摘要 · Abstract (English)

While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive processing before decisions can be made. In this domain, a sequential and interpretive understanding of visual cues is essential for effective perception, prediction, and planning. Nevertheless, common VQA benchmarks often focus on the accuracy of the final answer while overlooking the reasoning process that enables the generation of accurate responses. Moreover, existing methods lack a comprehensive framework for evaluating step-by-step reasoning in realistic driving scenarios. To address this gap, we propose DriveLMM-o1, a new dataset and benchmark specifically designed to advance step-wise visual reasoning for autonomous driving. Our benchmark features over 18k VQA examples in the training set and more than 4k in the test set, covering diverse questions on perception, prediction, and planning, each enriched with step-by-step reasoning to ensure logical inference in autonomous driving scenarios. We further introduce a large multimodal model that is fine-tuned on our reasoning dataset, demonstrating robust performance in complex driving scenarios. In addition, we benchmark various open-source and closed-source methods on our proposed dataset, systematically comparing their reasoning capabilities for autonomous driving tasks. Our model achieves a +7.49% gain in final answer accuracy, along with a 3.62% improvement in reasoning score over the previous best open-source model. Our framework, dataset, and model are available at https://github.com/ayesha-ishaq/DriveLMM-o1.

自动驾驶多步推理多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。