arXiv:2506.05667cs.CVcs.AI2025-06被引 18

首个面向VLA模型的驾驶决策基准,提升真实场景下动作预测能力。

DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models

  • 基于实车数据构建16,185组问答对,覆盖2,610个驾驶场景。
  • 引入驾驶员实际操作动作标签,准确率在无视觉或语言输入时分别下降3.3%和4.1%。
  • 树状评估框架支持细粒度分析,适合研究人机驾驶决策一致性。

视觉-语言-动作(VLA)模型推动了自动驾驶发展,但现有基准仍缺乏场景多样性、可靠的动作级标注及与人类偏好对齐的评估协议。为此,我们提出首个以动作为核心的基准DriveAction,涵盖从2,610个驾驶场景中生成的16,185组问答对。该基准利用自动驾驶车辆驾驶员主动采集的真实数据,确保场景广度与代表性;提供由驾驶员实际驾驶操作直接获取的高层离散动作标签;并采用以动作为中心的树状评估框架,明确关联视觉、语言与动作任务,支持全面与定向评估。实验表明,最先进的视觉-语言模型(VLMs)需同时依赖视觉与语言引导才能实现准确动作预测:平均而言,缺少视觉输入时准确率下降3.3%,缺少语言输入时下降4.1%,两者均缺失时下降8.0%。评估结果稳定可靠,可精准识别模型瓶颈,为实现更类人的自动驾驶决策提供新洞见与严谨基础。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protocols aligned with human preferences. To address these limitations, we introduce DriveAction, the first action-driven benchmark specifically designed for VLA models, comprising 16,185 QA pairs generated from 2,610 driving scenarios. DriveAction leverages real-world driving data proactively collected by drivers of autonomous vehicles to ensure broad and representative scenario coverage, offers high-level discrete action labels collected directly from drivers' actual driving operations, and implements an action-rooted tree-structured evaluation framework that explicitly links vision, language, and action tasks, supporting both comprehensive and task-specific assessment. Our experiments demonstrate that state-of-the-art vision-language models (VLMs) require both vision and language guidance for accurate action prediction: on average, accuracy drops by 3.3% without vision input, by 4.1% without language input, and by 8.0% without either. Our evaluation supports precise identification of model bottlenecks with robust and consistent results, thus providing new insights and a rigorous foundation for advancing human-like decisions in autonomous driving.

自动驾驶多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。