arXiv:2501.03225cs.CVcs.AI2025-01CVPR被引 42

自动将开放问答转为选择题,实现更客观的视觉语言模型评估

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

论文配图:Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
图 1 · 摘自论文原文
  • 构建智能代理框架AutoConverter,自动转换开放式问题为选择题
  • 生成9018道高质量选择题,模型在新基准上表现一致且略低于人工题
  • 适合需要可重复、标准化评估的VLM研究者使用

视觉语言模型(VLMs)的快速发展亟需严谨可靠的评估方法。当前视觉问答(VQA)基准多采用开放性问题,因自然语言回答的多样性,难以实现精准评估。为此,我们提出AutoConverter——一个智能代理框架,可自动将开放式问题转化为多项选择题,实现客观评估并大幅降低人工创建选择题的成本。实验表明,AutoConverter生成的问题既正确又具挑战性,且在这些题目上,VLMs的表现与人类创建的问题相比,准确率稳定或略低。基于此,我们构建了VMCBench,将20个现有VQA数据集统一转化为多项选择格式,共包含9,018道题目。我们在该基准上全面评估了33个先进VLMs,确立了可扩展、一致且可复现的VLM评估新标准。

原文摘要 · Abstract (English)

The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly multiple-choice question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation.

视觉语言模型自动评估选择题生成基准构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。