用百万级驾驶知识数据集测试大模型能否考过驾照。
Can LVLMs Obtain a Driver's License? A Benchmark Towards Reliable AGI for Autonomous Driving
- 构建包含百万条驾驶知识的数据集IDKB,覆盖理论到实操
- 15个大模型在该数据集上表现参差,揭示通用能力不足
- 适合关注自动驾驶安全与可解释性的研究者使用
大型视觉-语言模型(LVLMs)近年来受到广泛关注,许多研究致力于利用其通用知识提升自动驾驶模型的可解释性与鲁棒性。然而,现有LVLM通常依赖大规模通用数据集,缺乏专业驾驶所需的领域知识。现有视觉-语言驾驶数据集主要关注场景理解与决策,未提供交通规则与驾驶技能的明确指导,而这些正是驾驶安全的核心要素。为弥合这一差距,我们提出IDKB,一个涵盖超百万条数据项的大规模数据集,内容来自多个国家的驾驶手册、理论考试题及模拟路考数据。该数据集完整覆盖从理论到实践的全部显式驾驶知识,类似于考取驾照的过程。我们对15个主流LVLM进行了全面测试,并提供了详尽分析。此外,通过微调常用模型,实现了显著性能提升,进一步验证了该数据集的价值。项目页面详见: https://4dvlab.github.io/project_page/idkb.html
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have recently garnered significant attention, with many efforts aimed at harnessing their general knowledge to enhance the interpretability and robustness of autonomous driving models. However, LVLMs typically rely on large, general-purpose datasets and lack the specialized expertise required for professional and safe driving. Existing vision-language driving datasets focus primarily on scene understanding and decision-making, without providing explicit guidance on traffic rules and driving skills, which are critical aspects directly related to driving safety. To bridge this gap, we propose IDKB, a large-scale dataset containing over one million data items collected from various countries, including driving handbooks, theory test data, and simulated road test data. Much like the process of obtaining a driver's license, IDKB encompasses nearly all the explicit knowledge needed for driving from theory to practice. In particular, we conducted comprehensive tests on 15 LVLMs using IDKB to assess their reliability in the context of autonomous driving and provided extensive analysis. We also fine-tuned popular models, achieving notable performance improvements, which further validate the significance of our dataset. The project page can be found at: \url{https://4dvlab.github.io/project_page/idkb.html}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。