arXiv:2505.24120cs.CVcs.AI2025-05被引 5

构建中文科学推理多模态评测集,检验视觉语言模型的跨学科理解能力

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

  • 设计基于真实科学场景的图文问答对,融合领域知识与视觉证据分析
  • 15个VLM在该基准上最高仅49.6%准确率,暴露推理能力严重不足
  • 适用于评估模型在物理、化学等STEM领域的复杂推理表现

视觉语言模型(VLMs)在多模态理解方面取得显著进展,但其科学推理能力尚未得到充分评估。现有多模态基准主要聚焦通用图像理解或文本驱动推理,缺乏需要结合领域知识与视觉证据的真正科学情境。为此,我们提出CSVQA,一个面向科学推理的诊断性多模态基准,通过1,378个精心构建的跨学科问题-答案对,要求模型整合领域知识、分析视觉证据并进行高阶推理。相比以往基准,CSVQA更强调真实科学内容与复杂推理过程。我们还提出严谨的评估协议,基于专家标注的解释系统性检验模型预测是否具备有效中间推理步骤。对15个VLM的全面评估显示显著性能差异,即使顶尖专有模型也仅达49.6%准确率。该结果凸显提升VLM科学推理能力的紧迫性。CSVQA已公开发布于https://huggingface.co/datasets/Skywork/CSVQA。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remain inadequately assessed. Current multimodal benchmarks predominantly evaluate generic image comprehension or text-driven reasoning, lacking authentic scientific contexts that require domain-specific knowledge integration with visual evidence analysis. To fill this gap, we present CSVQA, a diagnostic multimodal benchmark specifically designed for evaluating scientific reasoning through domain-grounded visual question answering. Our benchmark features 1,378 carefully constructed question-answer pairs spanning diverse STEM disciplines, each demanding domain knowledge, integration of visual evidence, and higher-order reasoning. Compared to prior multimodal benchmarks, CSVQA places greater emphasis on real-world scientific content and complex reasoning. We additionally propose a rigorous evaluation protocol to systematically assess whether model predictions are substantiated by valid intermediate reasoning steps based on curated explanations. Our comprehensive evaluation of 15 VLMs on this benchmark reveals notable performance disparities, as even the top-ranked proprietary model attains only 49.6% accuracy. This empirical evidence underscores the pressing need for advancing scientific reasoning capabilities in VLMs. Our CSVQA is released at https://huggingface.co/datasets/Skywork/CSVQA

多模态科学推理中文数据集VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。