arXiv:2607.00159cs.CLcs.CV2026-07中稿 · ECCV

发现现有视觉问答评估存在系统性缺陷,提出修复与增强方案。

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

论文配图:Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
图 1 · 摘自论文原文
  • 系统审计发现答案不可推导、问题描述不清等漏洞
  • 修复后模型排名显著变化,证明原评测结果失真
  • 新增多实体场景以考验模型真实推理能力

基于知识的视觉问答(KB-VQA)旨在评估视觉语言模型是否能从外部结构化知识中检索、定位并推理。当前普遍以答案准确率为评价指标,隐含假设答案应可由知识库推导、问题需明确且视觉场景需引发语义歧义。然而本文审计发现,现有基准数据集普遍存在答案缺失或矛盾、问题表述不全、场景过于简单等问题,导致准确率无法反映真实推理能力。此外,多数样本仅涉及单一实体,无需复杂视觉-知识映射。即使在控制模型架构下,这些问题仍造成模型排名扭曲和能力误判。为此,本文提出:(1)原则性审计修复协议,恢复答案可推导性与问题清晰度;(2)受控多实体增强协议,引入视觉模糊性以挑战初始检索与定位推理。在修复与增强后的设置下,性能趋势明显改变。研究呼吁重新审视评估流程,并设计更注重可验证推理的交互式KB-VQA基准。

原文摘要 · Abstract (English)

Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external structured knowledge beyond visual evidence. In practice, answer accuracy is widely adopted as the primary evaluation metric, implicitly treating correctness as a proxy for knowledge-grounded reasoning. However, for existing KB-VQA benchmarks, this proxy relies on critical assumptions that are often overlooked and rendered unreliable by benchmark issues: annotated answer must be derivable from the associated knowledge base, question must be well-posed with sufficient constraints, and visual setting must meaningfully require grounded disambiguation. In this work, we show that these assumptions are systematically violated in existing KB-VQA benchmarks. Our audit reveals substantial instances with missing or contradicted answers and underspecified questions that render accuracy a misleading metric. Furthermore, we find that existing datasets rely on visually trivial, single-entity scenes that bypass the need for sophisticated visual-to-knowledge mapping. We demonstrate that even with controlled architectures, these flaws lead to distorted model rankings and overestimations of reasoning capabilities. To address this, we introduce (1) a principled audit-and-repair protocol that restores answer derivability and question clarity, and (2) a controlled multi-entity augmentation protocol that introduces visual ambiguity to challenge initial retrieval and grounded reasoning. Re-evaluation under corrected and augmented settings yields markedly different performance trends. Our findings call for rethinking evaluation protocols and designing more interaction-aware KB-VQA benchmarks that prioritize verifiable reasoning over simple matching.

视觉问答知识推理评测基准模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。