构建端到端自动驾驶评估基准,揭示视觉语言模型在感知到决策链中的失效模式。
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
- 分层设计感知到决策的渐进式评估流程,覆盖物体、场景和决策三类问题。
- 6650个问题中识别出逻辑推理错误与语义特征遗漏等关键失败模式。
- 提供可自动标注错误类型的轻量分析模型,适合安全敏感型AI系统研发者。
自动驾驶需要在复杂场景中实现可靠感知与安全决策。近年视觉语言模型(VLMs)展现出推理与泛化能力,为自动驾驶带来新可能;但现有基准常将感知与决策分开评估,依赖仅选答案的格式限制故障分析,或通过大模型评分生成长文本引入评价偏差。为此,我们提出Drive-P2D,一个包含6,650个问题的渐进式感知-决策评估基准,涵盖物体、场景和决策层级。该基准采用分离的推理与答案评分协议:最终答案客观打分,推理过程则用于识别感知-决策链条中暴露的错误模式。我们在主流VLM上评估所有及高风险场景表现,并通过相关性分析与相似场景鲁棒性测试刻画感知-决策能力边界。推理进一步揭示了逻辑推理错误与语义特征遗漏等失败模式,并训练轻量级分析模型实现大规模推理错误类型自动化标注。这些设计为构建更安全可靠的现实自动驾驶VLM提供了实用洞见。
原文摘要 · Abstract (English)
Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reasoning and generalization abilities, opening new possibilities for autonomous driving; however, existing benchmarks often evaluate perception and decision-making separately, limit failure analysis with choice-only formats, or introduce evaluation bias through LLM-scored long-form outputs. To address these issues, we present Drive-P2D, a progressive perception-to-decision benchmark with 6,650 questions across Object, Scene, and Decision levels. Drive-P2D adopts a separated reasoning-and-answer protocol: final answers are scored objectively, while reasoning is analyzed to identify error modes exposed along the progressive perception-to-decision chain. We evaluate mainstream VLMs across all and high-risk scenarios, and further characterize the perception-to-decision capability boundary through correlation analysis and similar-scene robustness testing. Reasoning further exposes failure modes such as logical reasoning errors and semantic feature omissions, and we train a lightweight analyzer model to automate large-scale error-mode annotation of reasoning. Together, these designs provide practical insights for building safer and more reliable VLMs for real-world autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。