首个面向城市自动驾驶的视觉语言模型多选题评测基准,提升评估可复现性。
AutoDrive-QA: A Multiple-Choice Benchmark for Vision-Language Evaluation in Urban Autonomous Driving
- 将开放问答数据集转为带真实错误干扰项的多选题
- GPT-4V零样本最高准确率达69.8%,微调后性能提升6个百分点
- 适合评估自动驾驶中感知、预测与规划任务的视觉语言模型
城市驾驶场景中视觉语言模型(VLMs)的评估仍具挑战性,因现有基准依赖开放式回答,存在歧义、标注成本高且评分不一致的问题。这阻碍了安全可靠城市智能交通AI的发展。我们提出AutoDrive-QA,首个将开放问答驾驶数据集(DriveLM、NuScenes-QA、LingoQA)系统转化为结构化多选题(MCQs)的基准,干扰项基于五大真实错误类别:驾驶领域误解、逻辑矛盾、传感器输入误读、计算疏漏及问题模糊。该框架支持在复杂城市场景中对感知、预测和规划任务进行可复现、可解释的VLM评估。实验表明,微调LLaVA-1.5-7B在各任务上准确率提升约6个百分点;GPT-4V实现最强零样本性能,最高准确率达69.8%;Qwen2-VL模型在多视角设置下也表现优异。此外,传统指标如BLEU和CIDEr无法区分强弱模型。AutoDrive-QA通过提供客观、领域相关的评估协议,推动城市AI系统的透明化测评,助力更安全可信的自动驾驶技术发展。
原文摘要 · Abstract (English)
Evaluating vision-language models (VLMs) in urban driving contexts remains challenging, as existing benchmarks rely on open-ended responses that are ambiguous, annotation-intensive, and inconsistent to score. This lack of standardized evaluation slows progress toward safe and reliable AI for urban mobility. We introduce AutoDrive-QA, the first benchmark that systematically converts open-ended driving QA datasets (DriveLM, NuScenes-QA, LingoQA) into structured multiple-choice questions (MCQs) with distractors grounded in five realistic error categories: Driving Domain Misconceptions, Logical Inconsistencies, Misinterpreted Sensor Inputs, Computational Oversights, and Question Ambiguity. This framework enables reproducible and interpretable evaluation of VLMs across perception, prediction, and planning tasks in complex urban scenes. Experiments show that fine-tuning LLaVA-1.5-7B improves accuracy by about six percentage points across tasks, GPT-4V achieves the strongest zero-shot performance with up to 69.8% accuracy, and Qwen2-VL models also perform competitively, particularly in multi-view settings. Moreover, traditional metrics such as BLEU and CIDEr fail to distinguish strong from weak models. By providing an objective, domain-grounded evaluation protocol, AutoDrive-QA contributes to more transparent benchmarking of urban AI systems, supporting the development of safer and more trustworthy autonomous driving technologies for smart cities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。