arXiv:2605.10187cs.CV2026-05ACL被引 2

构建跨学科多模态科学推理评测基准,揭示大模型在复杂推理中的短板。

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

论文配图:SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
图 1 · 摘自论文原文
  • 覆盖54个学科领域,融合方程、图表等专业视觉信息
  • 46%任务含专家标注解题路径,支持过程可追溯评估
  • 适合评估多步推理与跨学科知识融合能力的研究者

科学推理是人类智能的关键,需整合多模态输入、领域知识和多步推断。现有多模态大模型(MLLM)评测基准难以捕捉严谨评估所需的推理过程复杂性与可追溯性。为此,我们提出SciVQR,一个涵盖数学、物理、化学、地理、天文、生物等54个子领域的多模态基准。该数据集包含领域特定的视觉元素,如公式、图表与示意图,要求模型结合视觉理解进行推理。任务从基础事实回忆到复杂多步推断均有覆盖,其中46%的任务配有专家撰写的解题过程。SciVQR不仅评估最终答案,还分析推理路径,揭示模型决策逻辑。对主流MLLM(含闭源与开源)的评估显示,其在复杂多模态推理任务中存在显著局限,凸显了提升多步推理与跨学科知识融合能力的必要性。数据集与评估代码已公开于https://github.com/CASIA-IVA-Lab/SciVQR。

原文摘要 · Abstract (English)

Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference across various subjects. Existing benchmarks for multimodal large language models (MLLMs) often fail to capture the complexity and traceability of reasoning processes necessary for rigorous evaluation. To fill this gap, we introduce SciVQR, a multimodal benchmark covering 54 subfields in mathematics, physics, chemistry, geography, astronomy, and biology. SciVQR includes domain-specific visuals, such as equations, charts, and diagrams, and challenges models to combine visual comprehension with reasoning. The tasks range from basic factual recall to complex, multi-step inferences, with 46% including expert-authored solutions. SciVQR not only evaluates final answers but also examines the reasoning process, providing insights into how models reach their conclusions. Our evaluation of leading MLLMs, including both proprietary and open-source models, reveals significant limitations in handling complex multimodal reasoning tasks, underscoring the need for improved multi-step reasoning and better integration of interdisciplinary knowledge in advancing MLLMs toward true scientific intelligence. The dataset and evaluation code are publicly available at https://github.com/CASIA-IVA-Lab/SciVQR.

科学推理多模态评测跨学科大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。