arXiv:2502.14917cs.CVcs.AI2025-02被引 32

让AI像人一样思考开车,从看懂场景到控制车辆一气呵成。

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

  • 构建人类驾驶思维链的多模态大模型框架,融合局部视频与全局鸟瞰图。
  • 在CARLA基准上实现端到端驾驶最优性能,跨场景泛化能力显著提升。
  • 专为3D空间理解设计首个大规模驾驶指令VQA数据集,适合自动驾驶研究者。

端到端自动驾驶直接将原始传感器输入映射为低级车辆控制指令,是具身智能的重要组成部分。尽管多模态大语言模型(MLLM)在高层交通场景语义理解方面取得进展,但如何有效将这些概念性理解转化为低级运动控制指令,并实现跨场景的泛化与共识仍具挑战。本文提出Sce2DriveX,一种类人驾驶思维链(CoT)推理的MLLM框架。该框架通过联合学习局部场景视频与全局鸟瞰图(BEV),深入理解长时序时空关系与道路拓扑结构,增强对三维动态/静态场景的综合感知与推理能力,实现跨场景驾驶泛化。在此基础上,重构人类驾驶中隐含的认知链条,涵盖场景理解、元动作推理、行为解析、运动规划与控制,进一步缩小自动驾驶与人类思维之间的差距。为提升模型性能,我们构建了首个面向3D空间理解与长轴任务推理的大型视觉问答(VQA)驾驶指令数据集。大量实验表明,Sce2DriveX在从场景理解到端到端驾驶的全链路任务中均达到当前最优表现,并在CARLA Bench2Drive基准上展现出强鲁棒性与泛化能力。

原文摘要 · Abstract (English)

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-level traffic scene semantic understanding, it remains challenging to effectively translate these conceptual semantics understandings into low-level motion control commands and achieve generalization and consensus in cross-scene driving. We introduce Sce2DriveX, a human-like driving chain-of-thought (CoT) reasoning MLLM framework. Sce2DriveX utilizes multimodal joint learning from local scene videos and global BEV maps to deeply understand long-range spatiotemporal relationships and road topology, enhancing its comprehensive perception and reasoning capabilities in 3D dynamic/static scenes and achieving driving generalization across scenes. Building on this, it reconstructs the implicit cognitive chain inherent in human driving, covering scene understanding, meta-action reasoning, behavior interpretation analysis, motion planning and control, thereby further bridging the gap between autonomous driving and human thought processes. To elevate model performance, we have developed the first extensive Visual Question Answering (VQA) driving instruction dataset tailored for 3D spatial understanding and long-axis task reasoning. Extensive experiments demonstrate that Sce2DriveX achieves state-of-the-art performance from scene understanding to end-to-end driving, as well as robust generalization on the CARLA Bench2Drive benchmark.

自动驾驶多模态大模型思维链端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。