arXiv:2606.20980cs.CVcs.AI2026-06

对比人类与视觉语言模型在利马、纽约的自动驾驶场景表现

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

论文配图:Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
图 1 · 摘自论文原文
  • 用双城实录视频+VQA任务测试人类与VLMs应对新地理场景能力
  • 人类回答一致性高,地理来源不影响表现;VLMs响应差异大且受问题类型影响
  • 发现跨城市泛化中人类与VLMs行为差异显著,适合评估AI驾驶系统鲁棒性

随着自动驾驶汽车向国际扩张并采用多模态系统(如视觉语言模型)作为行动决策的认知核心,这些系统在新环境中,尤其是分布外(OOD)边缘场景下的泛化能力如何?本文通过全因子分析,对比了利马和纽约市的人类驾驶员以及视觉语言模型,使用两地采集的行车记录视频,在视觉问答(VQA)范式下提出四类问题:事实类、评价类、反事实类和推理类。选取这两座目前无自动驾驶公司运营的高挑战性城市,结果表明人类回答具有一致性,不受地理来源影响;而视觉语言模型的表现则受问题类型显著调节,且未发现明显地理差异,可能源于场景高度分布外。数据集已公开于:https://huggingface.co/datasets/Artificio/robusto-2

原文摘要 · Abstract (English)

As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://huggingface.co/datasets/Artificio/robusto-2

自动驾驶视觉语言模型泛化能力多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。