arXiv:2504.04348cs.CV2025-04CVPR被引 193

构建支持反事实推理的3D自动驾驶视觉语言数据集

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

  • 通过反事实推理生成大规模高质量合成标注数据
  • 在DriveLM问答与nuScenes规划任务上显著提升性能
  • 适合研究多模态决策与自动驾驶智能体设计的学者

视觉语言模型的发展推动了其在自动驾驶中的应用,但将能力从2D扩展到全3D理解对实际应用至关重要。为此,我们提出OmniDrive,一个通过反事实推理对齐智能体模型与3D驾驶任务的综合性视觉语言数据集。该方法通过评估潜在情景及其结果来增强决策能力,类比人类驾驶员考虑替代行为。基于反事实的合成数据标注流程生成大规模、高质量数据,提供更密集的监督信号,弥合规划轨迹与语言推理之间的差距。此外,我们探索了两种先进的OmniDrive-Agent框架——Omni-L和Omni-Q,以评估视觉语言对齐与3D感知的重要性,揭示了设计有效大模型智能体的关键洞见。在DriveLM Q&A基准和nuScenes开环规划任务上均取得显著提升,验证了数据集与方法的有效性。

原文摘要 · Abstract (English)

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for real-world applications. To address this challenge, we propose OmniDrive, a holistic vision-language dataset that aligns agent models with 3D driving tasks through counterfactual reasoning. This approach enhances decision-making by evaluating potential scenarios and their outcomes, similar to human drivers considering alternative actions. Our counterfactual-based synthetic data annotation process generates large-scale, high-quality datasets, providing denser supervision signals that bridge planning trajectories and language-based reasoning. Futher, we explore two advanced OmniDrive-Agent frameworks, namely Omni-L and Omni-Q, to assess the importance of vision-language alignment versus 3D perception, revealing critical insights into designing effective LLM-agents. Significant improvements on the DriveLM Q\&A benchmark and nuScenes open-loop planning demonstrate the effectiveness of our dataset and methods.

自动驾驶视觉语言反事实推理3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。