arXiv:2503.10627cs.CVcs.AI2025-03ACL被引 23

评测大模型在科学问题上的知识理解与视觉推理能力

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

  • 构建五种版本测试题,区分知识需求与图文信息依赖程度
  • 大模型在高知识需求任务中准确率不足40%,视觉推理能力弱
  • 适合关注多模态模型科学推理能力的研究者使用

大型多模态模型(LMMs)在科学问题求解中应用迅速发展,但其精细能力仍不明确。本文提出SciVerse,一个包含5,735个测试实例的多模态科学评估基准,涵盖五个不同版本。通过设计知识零、轻、丰富三类题目,探究模型对科学知识的理解;通过视觉丰富和仅视觉版本,分析模型对图表信息的解读能力。同时,提出一种新的科学链式思维(CoT)评估策略,分步检测输出中的知识与逻辑错误。在SciVerse上对多种LMM的广泛评估揭示了其在科学专业性上的显著局限,为未来发展方向提供新洞见。

原文摘要 · Abstract (English)

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In this paper, we introduce SciVerse, a multi-modal scientific evaluation benchmark to thoroughly assess LMMs across 5,735 test instances in five distinct versions. We aim to investigate three key dimensions of LMMs: scientific knowledge comprehension, multi-modal content interpretation, and Chain-of-Thought (CoT) reasoning. To unveil whether LMMs possess sufficient scientific expertise, we first transform each problem into three versions containing different levels of knowledge required for solving, i.e., Knowledge-free, -lite, and -rich. Then, to explore how LMMs interpret multi-modal scientific content, we annotate another two versions, i.e., Vision-rich and -only, marking more question information from texts to diagrams. Comparing the results of different versions, SciVerse systematically examines the professional knowledge stock and visual perception skills of LMMs in scientific domains. In addition, to rigorously assess CoT reasoning, we propose a new scientific CoT evaluation strategy, conducting a step-wise assessment on knowledge and logical errors in model outputs. Our extensive evaluation of different LMMs on SciVerse reveals critical limitations in their scientific proficiency and provides new insights into future developments. Project page: https://sciverse-cuhk.github.io

多模态模型科学推理视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。