构建细粒度数据集,评估视觉语言模型在自动驾驶中的推理能力
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
- 设计细粒度闭合问答数据集,覆盖5大领域29项任务
- 发现通用模型在动态场景推理上存在明显短板
- 适合研究自动驾驶多模态理解与认知推理的学者
现有自动驾驶视觉语言模型评估基准主要依赖开放形式的视觉问答任务,难以全面检验复杂驾驶场景下的理解能力。为此,本文提出VLADBench,一个包含闭合问答的细粒度数据集,任务从静态基础知识逐步过渡到动态道路场景的高级推理。该数据集涵盖交通知识理解、通用元素识别、交通图生成、目标属性理解及自身决策规划共5个核心领域,细分为11个二级维度和29个三级任务。对通用模型与特定领域(DS)模型在该基准上的全面评估揭示了其在自动驾驶场景下的优势与关键局限。通过基于140万条公开来源的特定领域问答数据,从小规模模型出发训练各领域专用模型,实验表明该基准为更全面评估视觉语言模型在自动驾驶中的认知与推理能力提供了关键支撑,推动发展更具认知深度与推理能力的自动驾驶系统。
原文摘要 · Abstract (English)
Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce $\textbf{VLADBench}$, a challenging and fine-grained dataset featuring close-form QAs that progress from static foundational knowledge and elements to advanced reasoning for dynamic on-road situations. The elaborate $\textbf{VLADBench}$ spans 5 key domains: Traffic Knowledge Understanding, General Element Recognition, Traffic Graph Generation, Target Attribute Comprehension, and Ego Decision-Making and Planning. These domains are further broken down into 11 secondary aspects and 29 tertiary tasks for a granular evaluation. A thorough assessment of general and domain-specific (DS) VLMs on this benchmark reveals both their strengths and critical limitations in AD contexts. To further exploit the cognitive and reasoning interactions among the 5 domains for AD understanding, we start from a small-scale VLM and train the DS models on individual domain datasets (collected from 1.4M DS QAs across public sources). The experimental results demonstrate that the proposed benchmark provides a crucial step toward a more comprehensive assessment of VLMs in AD, paving the way for the development of more cognitively sophisticated and reasoning-capable AD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。