arXiv:2409.13730cs.AIcs.CL2024-09被引 14

构建K12科学多模态推理评测基准,覆盖数理化三大学科。

VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning

  • 构建涵盖数理化三学科的3000道题多模态评测集。
  • 模型最高准确率:数学53.4%(Claude3.5-Sonnet),物理38.2%(GPT-4o)。
  • 揭示闭源模型整体优于开源模型,推动科学推理能力提升。

多模态大语言模型(MLLMs)通过融合文本与视觉信息,在复杂场景中展现出色的视觉理解能力。尽管已有多个基准用于评估MLLM在视觉问答到复杂问题求解任务中的表现,但多数集中于数学或通用视觉理解,忽视了物理、化学等关键科学领域。为此,我们精心构建了名为VisScience的综合性评测基准,用于评估数学、物理和化学三大学科中的多模态科学推理能力。该基准包含3000道来自K12教育阶段(小学至高中)的题目,每学科各1000道,覆盖21个细分主题,按五级难度分级。基于此,我们对25个代表性MLLM进行了详细评估。实验结果表明,闭源模型整体表现更优,最佳成绩为:数学53.4%(Claude3.5-Sonnet)、物理38.2%(GPT-4o)、化学47.0%(Gemini-1.5-Pro)。这些结果揭示了当前模型在科学推理中的优势与局限,为未来改进提供方向。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have demonstrated promising capabilities across various tasks by integrating textual and visual information to achieve visual understanding in complex scenarios. Despite the availability of several benchmarks aims to evaluating MLLMs in tasks from visual question answering to complex problem-solving, most focus predominantly on mathematics or general visual understanding tasks. This reveals a critical gap in current benchmarks, which often overlook the inclusion of other key scientific disciplines such as physics and chemistry. To address this gap, we meticulously construct a comprehensive benchmark, named VisScience, which is utilized to assess the multi-modal scientific reasoning across the three disciplines of mathematics, physics, and chemistry. This benchmark comprises 3,000 questions drawn from K12 education - spanning elementary school through high school - equally distributed across three disciplines, with 1,000 questions per discipline. The questions within VisScience span 21 distinct subjects and are categorized into five difficulty levels, offering a broad spectrum of topics within each discipline. With VisScience, we present a detailed evaluation of the performance of 25 representative MLLMs in scientific reasoning. Experimental results demonstrate that closed-source MLLMs generally outperform open-source models. The best performance observed include a 53.4\% accuracy in mathematics by Claude3.5-Sonnet, 38.2\% in physics by GPT-4o, and 47.0\% in chemistry by Gemini-1.5-Pro. These results underscore the strengths and limitations of MLLMs, suggesting areas for future improvement and highlighting the importance of developing models that can effectively handle the diverse demands of multi-modal scientific reasoning.

多模态科学推理教育评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。