arXiv:2606.09644cs.CLcs.CV2026-06被引 1

测试多视角自动驾驶模型是否看对了位置。

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

论文配图:Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving
图 1 · 摘自论文原文
  • 设计新基准,让模型选出问题对应的正确摄像头视角
  • 在73个场景中验证,模型常选错视角但答对题
  • 适合评估自动驾驶视觉决策的可靠性

多模态大语言模型在视觉推理任务中表现优异,但仅靠答案正确性无法判断其是否基于正确视觉证据。这一问题在自动驾驶的多视角场景中尤为关键,因为模型可能用错误摄像头的画面生成合理回答。我们构建了一个多视角视觉问答基准,使用六个同步的NuScenes视角和问题,要求模型识别支持答案的相机视图并作答。该基准包含122个以冲突为核心的问答对,来自73个场景,涵盖因果推理、反事实推理和意图预测。视图标签由自动冲突挖掘流程生成,并经人工标注确认。评估三种设置:相机视图选择、已知最优视图的最优问答、联合预测(模型一次完成选视图和作答)。答案采用多项选择与自由形式两种格式,结构化答案用精确匹配评估,自由形式由LLM裁判打分。通过将视觉来源识别与答案正确性分离,该基准揭示了仅看答案会遗漏的错误定位问题。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a model relied on the correct visual evidence. This gap is particularly important in multi-view driving scenes used for autonomous driving, where a model can produce a plausible answer while grounding it in the wrong camera view. We introduce a multi-view visual question answering benchmark for evaluating evidence-source identification: given six synchronized NuScenes views and a question, the model must identify the supporting camera view and answer the question. The benchmark contains 122 conflict-centric question-answer pairs from 73 scenes, spanning causality, counterfactual reasoning, and intent prediction. View labels are proposed by an automatic conflict-mining pipeline and manually verified by annotators. We evaluate three settings: camera-view selection, oracle QA given the golden view, and joint prediction in which the model selects a view and answers in one pass. Answers are evaluated in both multiple-choice and free-form formats, using exact match for structured predictions and an LLM judge for free-form responses. By explicitly separating visual-source identification from answer correctness, the benchmark exposes grounding failures that answer-only evaluation misses.

多视角视觉推理自动驾驶模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。