首个评估多模态大模型事实准确性的综合基准
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
- 按客观事件与常识分9类任务,覆盖9个主题
- 18个主流多模态模型在该基准上表现参差不齐
- 适合研究模型幻觉、评测可靠性的研究人员
多模态大语言模型(MLLMs)在各领域的广泛应用凸显了其输出可靠性与准确性的重要性,尤其体现在基于事实信息(如通用和领域知识)生成内容的能力。本文提出SimpleVQA,首个全面评估MLLM回答自然语言短问题时事实性能力的多模态基准。该基准具备六大特点:涵盖多种任务与场景、保证高质量且具有挑战性的问题、采用静态永恒的参考答案、评估过程简单。我们把视觉问答任务分为围绕客观事件或常识的9类,并置于9个主题中。通过严格的质量控制流程,确保答案简洁清晰,便于使用大模型作为评判者进行低方差评估。利用SimpleVQA,我们对18个领先多模态大模型和8个纯文本大模型进行了全面评估,深入分析其图像理解与文本生成能力,并识别与剖析错误案例。
原文摘要 · Abstract (English)
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work, we introduce SimpleVQA, the first comprehensive multi-modal benchmark to evaluate the factuality ability of MLLMs to answer natural language short questions. SimpleVQA is characterized by six key features: it covers multiple tasks and multiple scenarios, ensures high quality and challenging queries, maintains static and timeless reference answers, and is straightforward to evaluate. Our approach involves categorizing visual question-answering items into 9 different tasks around objective events or common knowledge and situating these within 9 topics. Rigorous quality control processes are implemented to guarantee high-quality, concise, and clear answers, facilitating evaluation with minimal variance via an LLM-as-a-judge scoring system. Using SimpleVQA, we perform a comprehensive assessment of leading 18 MLLMs and 8 text-only LLMs, delving into their image comprehension and text generation abilities by identifying and analyzing error cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。