评测视觉语言模型在不同国家驾驶规则下的推理能力,发现其泛化能力不足。
GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving

- 构建六国5053个带验证的多选题,测试模型对本地交通规则的理解。
- 九种主流模型在跨区域任务中表现差异显著,最高差距达42%。
- 提出知识蒸馏方法,提升模型对地域驾驶政策的适应性,适合自动驾驶研发者使用。
面向自动驾驶的视觉语言模型虽具潜力,但对区域特定交通规则的处理能力仍待探索,影响其在全球部署的可靠性。为此,我们提出GeoDrive-Bench,一个涵盖六个不同国家、包含5,053个经人工验证的多选问答对的新基准,系统评估模型在感知、预测、规划和区域推理四类任务中的地理文化驱动推理能力。每道题目需模型根据视觉证据与本地交通惯例推断正确驾驶行为,且不提供国家标签。除评估外,我们设计一种知识蒸馏算法,将区域交通规则知识注入模型内部表征,使其更准确地结合场景理解与本地驾驶规范。在九种先进视觉语言模型上的实验表明,各模型在不同驾驶文化下的表现存在显著差异(最高相差42%),而本研究提出的基线模型在跨区域推理中表现更优。结果表明当前模型仍缺乏稳健的区域感知驾驶智能,凸显GeoDrive-Bench作为诊断与训练工具的价值。
原文摘要 · Abstract (English)
Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains underexplored, raising uncertainties about their deployment across diverse global settings. We therefore introduce GeoDrive-Bench, a novel benchmark that enables the systematic investigation of VLMs' geo-culturally grounded driving reasoning. We curated 5,053 human-validated multiple-choice QA pairs across six countries covering diverse driving cultures. Specifically, we emphasize four driving tasks: perception, prediction, planning, and region reasoning. Each question requires models to infer the correct driving behavior from visual evidence and local traffic conventions without explicit country labels. Beyond evaluation, we further design a distillation algorithm that injects region-specific traffic-rule knowledge into the internal representations of VLMs, enabling models to better align visual scene understanding with local driving policies. Experiments on nine state-of-the-art VLMs show substantial performance variations across geo-driving cultures for each task, while our proposed baseline models exhibit improved geo-cultural reasoning across regions. These results suggest that current VLMs still lack robust region-aware driving intelligence and highlight GeoDrive-Bench as a diagnostic and training-oriented testbed for deployable autonomous driving foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。