小模型在远距离交通感知上表现差,暴露了自动驾驶视觉短板
Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- 构建带距离标注的交通感知问答数据集,专注评估感知能力
- 最先进小模型在远距任务上准确率仅60%,远低于人类85%水平
- 特别难区分左右方向,适合关注模型鲁棒性的自动驾驶研究者
视觉语言模型(VLM)在多模态理解任务中表现强劲,具备泛化能力,是自动驾驶系统的潜在组件。但其在安全关键场景中需可靠感知能力,尤其对远距离物体(30+米)的识别。为此,我们提出首个聚焦交通场景感知问题的带距离标注的视觉问答基准DTPQA,排除推理类问题以确保评估仅反映感知性能。鉴于自动驾驶硬件算力有限,研究集中于小型VLMs。评估显示,尽管问题简单,当前最佳小模型平均准确率仅60%,远低于人类约85%的表现。值得注意的是,人类样本量较小,存在统计局限性。此外,区分左右方向等特定任务仍极具挑战。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are becoming increasingly powerful, demonstrating strong performance on a variety of tasks that require both visual and textual understanding. Their strong generalisation abilities make them a promising component for automated driving systems, which must handle unexpected corner cases. However, to be trusted in such safety-critical applications, a model must first possess a reliable perception system. Moreover, since critical objects and agents in traffic scenes are often at a distance, we require systems that are not "shortsighted", i.e., systems with strong perception capabilities at both close (up to 20 meters) and long (30+ meters) range. With this in mind, we introduce Distance-Annotated Traffic Perception Question Answering (DTPQA), the first Visual Question Answering (VQA) benchmark focused solely on perception-based questions in traffic scenes, enriched with distance annotations. By excluding questions that require reasoning, we ensure that model performance reflects perception capabilities alone. Since automated driving hardware has limited processing power and cannot support large VLMs, our study centers on smaller VLMs. More specifically, we evaluate several state-of-the-art (SOTA) small VLMs on DTPQA and show that, despite the simplicity of the questions, these models significantly underperform compared to humans (~60% average accuracy for the best-performing small VLM versus ~85% human performance). However, it is important to note that the human sample size was relatively small, which imposes statistical limitations. We also identify specific perception tasks, such as distinguishing left from right, that remain particularly challenging for these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。