arXiv:2508.02645cs.CV2025-08ICCV被引 2

发现视觉问答评估中存在显著波动,建议改用更可靠的评测方法。

Evaluating Variance in Visual Question Answering Benchmarks

  • 系统分析14个VQA基准的性能波动来源
  • 训练种子和指令微调导致结果差异显著
  • 提出填空式评测以降低随机性,提升可靠性

多模态大语言模型(MLLM)在视觉问答(VQA)任务中表现出强大推理与上下文理解能力。然而,现有评估多依赖点估计,忽视了由模型输出随机性、训练种子敏感性及超参数配置等因素引起的显著性能波动。本文分析了14个常用VQA基准中的变异性,涵盖视觉推理、文本理解与常识推理等多样化任务。系统研究了训练种子、框架非确定性、模型规模及扩展指令微调对性能变异的影响。同时探讨了填空式(Cloze-style)评估作为替代方案的有效性,发现其能有效降低随机性并提升各基准上的评估可靠性。研究揭示当前评估方法的局限性,呼吁采用考虑方差的评价范式,以推动MLLM更稳健、可靠的开发。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have emerged as powerful tools for visual question answering (VQA), enabling reasoning and contextual understanding across visual and textual modalities. Despite their advancements, the evaluation of MLLMs on VQA benchmarks often relies on point estimates, overlooking the significant variance in performance caused by factors such as stochastic model outputs, training seed sensitivity, and hyperparameter configurations. This paper critically examines these issues by analyzing variance across 14 widely used VQA benchmarks, covering diverse tasks such as visual reasoning, text understanding, and commonsense reasoning. We systematically study the impact of training seed, framework non-determinism, model scale, and extended instruction finetuning on performance variability. Additionally, we explore Cloze-style evaluation as an alternate assessment strategy, studying its effectiveness in reducing stochasticity and improving reliability across benchmarks. Our findings highlight the limitations of current evaluation practices and advocate for variance-aware methodologies to foster more robust and reliable development of MLLMs.

视觉问答评估方法模型稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。