arXiv:2511.22341cs.CVcs.LG2025-11被引 1

发现多选VQA评估中隐藏的提示格式偏差,影响模型评测可靠性

Unexplored flaws in multiple-choice VQA evaluations

  • 测试7个大模型在48种提示格式下的表现,识别出隐蔽的格式偏差
  • 即使语义不变,微小提示变化也会导致答案偏差,影响率达20%以上
  • 现有缓解策略无效,适合关注评测公平性的研究者参考

多模态大语言模型(MLLM)在处理图文输入方面表现出强大能力。评估其能力的常见方法是多选视觉问答(VQA)。早期研究已揭示这些基准对答案顺序敏感,可通过精心设计缓解。然而,我们指出了提示格式中尚未被探索的其他偏差,质疑当前MLLM评估的可靠性。具体而言,通过涵盖7个MLLM和5个VQA数据集的大规模实验,分析了48种不同的提示格式变体,识别出三个关键的提示格式变化因素。研究发现,多选VQA对提示格式的微小变化极为敏感,即便这些变化在语义上是中立的。此外,这些偏差独立于已知的答案顺序偏差或模型对正确答案的信心而存在。最后,我们证明现有偏差缓解策略无法解决这些新发现的偏差。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in handling image-text inputs. A common way to assess this ability is through multiple-choice Visual Question Answering (VQA). Earlier works have already revealed that these benchmarks are sensitive to answer choice order, a limitation that can be mitigated through careful design. Yet, we highlight additional, unexplored biases in prompt formatting that question the reliability of current MLLM evaluations. Specifically, we identify three key variation factors in prompt formatting and analyze their impact through a large-scale study involving $\mathbf{\text{seven}}$ MLLMs and $\mathbf{\text{five}}$ VQA datasets, spanning $\mathbf{48}$ distinct $\mathbf{\text{prompt format variations}}$. Our findings reveal that multiple-choice VQA is highly sensitive to minor prompt format changes, even when these changes are semantically neutral. We further demonstrate that these biases persist independently of known order biases or the MLLM's confidence in the correct answer. Finally, we demonstrate that existing bias mitigation strategies fail to address these newly identified biases.

VQA评估提示偏差多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。