短答案VQA分数常被误读,因评分器对表达形式敏感而失真。
What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks

- 用人工验证的语义判断器审计3.7万条错误,发现评分器常误判正确答案。
- 文本丰富任务中,一半错误源于表达形式不符,非语义错误。
- 提取类和多段答案比数值答案更易受评分器影响,适合评估者参考。
短答案视觉问答基准将模型答案的语义正确性与评分器期望的表面形式混淆。本文通过人工验证的语义评判器(97.6%精确率)审计了超过3.7万条官方错误,结果表明:在文本丰富的基准上,高达一半的错误是语义正确但表面形式不匹配的答案被误判;另一纯文本评判器复现相同假阴性模式,证明该现象非单一模型偏差。不同答案类型受影响程度不同:抽取型和多段落答案比标量答案更易受评分器干扰。微小提示词或上下文改写即可显著改变评分结果,且不影响任务本质。通过仅依赖CPU的确定性合同修复验证,部分误判可恢复。研究建议:官方分数应附带语义审计与答案类型诊断以保持可解释性。
原文摘要 · Abstract (English)
Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, using a human-validated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。