arXiv:2605.20676cs.CV2026-05被引 1

评测视觉问答中答案与像素证据的联合准确性,揭示模型推理透明性短板。

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

论文配图:VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence
图 1 · 摘自论文原文
  • 构建多任务、多领域数据集,要求模型同时输出正确答案和精准分割掩码。
  • 引入新评估指标GROVE,强制兼顾文本准确性和视觉定位质量。
  • 发现当前最强模型在证据对齐上仍表现有限,凸显推理与视觉支撑的脱节。

建立模型预测与支持其结论的视觉证据之间的明确联系,对多模态推理的透明性和可靠性至关重要,但现有多模态大模型(MLLM)评估并未显式要求这种对齐。现有基准仅独立评估文本答案正确性或像素级定位,导致推理与定位耦合问题未被解决。我们提出VISTAQA,一个综合性基准,用于联合评估自由形式答案正确性与像素级证据定位。VISTAQA包含1,157个专家标注样本,覆盖六种任务类型和六种视觉领域,从直接感知到组合与关系推理。该基准要求模型不仅回答正确,还需提供支持答案的精确分割掩码,并包含无有效视觉证据的幻觉敏感样本。为支持此评估,我们引入GROVE,一种统一评估指标,通过样本级几何平均结合文本准确率与定位质量,确保任一维度无法弥补另一维度缺陷。在多种定位感知模型与通用MLLM混合流水线上的全面实验表明,即使最强系统在GROVE下表现也有限,凸显答案准确率与视觉证据对齐间的显著差距。

原文摘要 · Abstract (English)

Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VISTAQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VISTAQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VISTAQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.

视觉问答证据定位多模态评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。