arXiv:2508.21824cs.CV2025-08ICCV被引 9

测试大模型驾驶知识掌握程度,发现其在复杂场景下表现不足。

DriveQA: Passing the Driving Knowledge Test

  • 构建涵盖交通规则的图文基准数据集DriveQA,覆盖真实罕见边缘案例。
  • 现有大模型在数值推理与复杂路权场景中准确率显著下降,低于60%。
  • 微调和预训练可提升模型对标志识别与交叉口决策能力,适合自动驾驶研究者使用。

若大型语言模型(LLM)参加当前驾驶知识考试,能否通过?除主流自动驾驶基准中的空间与视觉问答任务外,驾驶知识测试需全面理解交通规则、标识及路权原则。人类驾驶员须识别真实数据集中罕见的边缘情况。本文提出DriveQA,一个大规模开源文本与视觉融合的基准,系统覆盖各类交通法规与场景。实验表明:(1)当前先进LLM与多模态大模型(MLLM)在基础规则上表现良好,但在数值推理、复杂路权场景、交通标志变化及空间布局理解上存在明显短板;(2)在DriveQA上微调可显著提升多个类别准确率,尤其在标志识别与交叉口决策方面;(3)通过控制变量的DriveQA-V分析了光照、视角、距离、天气等因素对模型性能的影响;(4)在DriveQA上预训练能提升下游驾驶任务表现,在nuScenes与BDD等真实数据集上取得更好结果,且模型能内化文本与合成交通知识,有效泛化至其他问答任务。

原文摘要 · Abstract (English)

If a Large Language Model (LLM) were to take a driving knowledge test today, would it pass? Beyond standard spatial and visual question-answering (QA) tasks on current autonomous driving benchmarks, driving knowledge tests require a complete understanding of all traffic rules, signage, and right-of-way principles. To pass this test, human drivers must discern various edge cases that rarely appear in real-world datasets. In this work, we present DriveQA, an extensive open-source text and vision-based benchmark that exhaustively covers traffic regulations and scenarios. Through our experiments using DriveQA, we show that (1) state-of-the-art LLMs and Multimodal LLMs (MLLMs) perform well on basic traffic rules but exhibit significant weaknesses in numerical reasoning and complex right-of-way scenarios, traffic sign variations, and spatial layouts, (2) fine-tuning on DriveQA improves accuracy across multiple categories, particularly in regulatory sign recognition and intersection decision-making, (3) controlled variations in DriveQA-V provide insights into model sensitivity to environmental factors such as lighting, perspective, distance, and weather conditions, and (4) pretraining on DriveQA enhances downstream driving task performance, leading to improved results on real-world datasets such as nuScenes and BDD, while also demonstrating that models can internalize text and synthetic traffic knowledge to generalize effectively across downstream QA tasks.

驾驶认知大模型评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。