首个中文视觉问答事实性评测基准,揭示大模型在看图识物与知识发现上的能力短板。
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
- 构建中文多跳问答数据集,区分图像识别与知识推理两个环节
- 覆盖8大主题56子类,1200+高质量短答案题,评估34个主流模型表现
- 开源代码与数据,适合中文多模态研究者、评测人员使用
大型视觉语言模型(LVLMs)的事实性评估滞后于其快速发展,难以真实反映模型的知识能力和可靠性。本文提出首个基于事实性的中文视觉问答评测基准——ChineseSimpleVQA,用于评估模型在8个主要领域和56个子领域的视觉事实性表现。该基准具有中文语境、多样知识类型、多跳问题设计、高质量数据、静态一致性及短答案可评性等特点。我们还构建了严谨的数据制作流程,并将视觉事实性解耦为‘看世界’(物体识别)和‘探知识’两个部分,便于分析模型的能力边界与执行机制。随后对34个先进开源与闭源模型进行了评估,揭示了该领域存在的关键性能差距。相关评估代码与数据已公开。
原文摘要 · Abstract (English)
The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in Chinese, named ChineseSimpleVQA, aimed at assessing the visual factuality of LVLMs across 8 major topics and 56 subtopics. The key features of this benchmark include a focus on the Chinese language, diverse knowledge types, a multi-hop question construction, high-quality data, static consistency, and easy-to-evaluate through short answers. Moreover, we contribute a rigorous data construction pipeline and decouple the visual factuality into two parts: seeing the world (i.e., object recognition) and discovering knowledge. This decoupling allows us to analyze the capability boundaries and execution mechanisms of LVLMs. Subsequently, we evaluate 34 advanced open-source and closed-source models, revealing critical performance gaps within this field. Our evaluation-friendly code and data have already been open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。