对比人类与视觉语言模型在秘鲁驾驶场景下的问答表现,发现认知对齐程度因问题类型而异。
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru
- 用秘鲁行车视频构建新数据集,测试模型在分布外场景的推理能力。
- 通过代表相似性分析发现,模型与人类在不同问题上的回答一致性差异显著。
- 适用于关注自动驾驶认知对齐、多模态模型泛化能力的研究者。
随着多模态基础模型开始在自动驾驶汽车中进行实验性部署,一个关键问题是:这些系统在特定驾驶情境下,尤其是分布外(out-of-distribution)的情况下,其响应与人类有多相似?为此,我们构建了Robusto-1数据集,采用来自秘鲁的行车记录仪视频——该国以驾驶行为激烈、交通指数高、街景中奇异物体比例高而著称,且这些物体很可能未出现在训练数据中。为初步从认知层面评估基础视觉语言模型(VLMs)与人类在驾驶中的表现差异,我们摒弃了边界框、分割图、占据图或轨迹估计等传统任务,转而采用多模态视觉问答(VQA),并运用系统神经科学中流行的代表相似性分析(RSA)方法,比较人类与机器的回答。结果显示,根据问题类型的不同,VLMs与人类的响应在某些情况下趋于一致,而在另一些情况下则显著偏离,揭示出两者间认知对齐程度存在明显差异。
原文摘要 · Abstract (English)
As multimodal foundational models start being deployed experimentally in Self-Driving cars, a reasonable question we ask ourselves is how similar to humans do these systems respond in certain driving situations -- especially those that are out-of-distribution? To study this, we create the Robusto-1 dataset that uses dashcam video data from Peru, a country with one of the worst (aggressive) drivers in the world, a high traffic index, and a high ratio of bizarre to non-bizarre street objects likely never seen in training. In particular, to preliminarly test at a cognitive level how well Foundational Visual Language Models (VLMs) compare to Humans in Driving, we move away from bounding boxes, segmentation maps, occupancy maps or trajectory estimation to multi-modal Visual Question Answering (VQA) comparing both humans and machines through a popular method in systems neuroscience known as Representational Similarity Analysis (RSA). Depending on the type of questions we ask and the answers these systems give, we will show in what cases do VLMs and Humans converge or diverge allowing us to probe on their cognitive alignment. We find that the degree of alignment varies significantly depending on the type of questions asked to each type of system (Humans vs VLMs), highlighting a gap in their alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。