arXiv:2607.08745cs.AIcs.CV2026-07

构建行车记录仪事故理解基准,评估模型安全推理能力

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

论文配图:AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding
图 1 · 摘自论文原文
  • 以真实驾驶事故为背景设计结构化问答任务
  • 覆盖天气、路况、标志、碰撞位置等10类安全相关维度
  • 适合评估自动驾驶系统在复杂场景下的可靠性

视觉语言模型、大语言模型及多模态大语言模型的进展提升了自动驾驶中的场景理解、决策与轨迹预测能力,但在安全关键事件上的推理可靠性仍难评估。为此,我们提出AUTOPILOT-VQA——一个面向行车记录仪视频理解的事故中心型视觉问答基准。该数据集通过围绕真实驾驶事故与近事故设计的结构化问题,评估系统对情境属性与事件细节的理解能力。涵盖天气光照、交通环境、道路布局、路面状况、标识、涉事主体、事故发生、撞击位置及可避免性等10类安全相关类别。通过要求模型回答关于时空上下文与事件层面细节的依据性问题,该基准推动从物体识别迈向时序关联、安全意识驱动的推理。数据集将作为AUTOPILOT CVPR 2026竞赛的一部分发布,提供标准化评测框架,助力发展更可解释、鲁棒且注重安全的视觉语言系统。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.

视觉问答自动驾驶安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。