arXiv:2506.10521cs.AIcs.CL2025-06NeurIPS被引 33

首个科学认知评测基准,测试AI在感知、理解与推理上的真实科研能力。

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

  • 构建三层次科学认知评测:感知信号、理解属性、比较推理。
  • 830个专家验证的跨学科多模态问答,覆盖66项任务,仅顶尖模型达34%准确率。
  • 适合科研型AI研发者,推动科学发现中的AI应用落地。

科学发现日益依赖基于信息密集型数据和领域专长的复杂多模态推理。得益于专家级科学基准,科学多模态大语言模型(MLLMs)有望显著提升真实工作流中的发现效率。然而,现有科学基准主要评估知识理解能力,对感知与推理能力评估不足。为此,我们提出科学家首考(SFE)基准,通过三个相互关联的层面——科学信号感知、科学属性理解、科学比较推理——评估MLLMs的科学认知能力。SFE包含830个专家验证的视觉问答对,涵盖三种题型,覆盖五个高价值学科的66项多模态任务。大量实验显示,当前最先进的GPT-o3和InternVL-3在SFE上分别仅达到34.08%和26.52%的准确率,表明其在科学领域仍有巨大提升空间。我们希望SFE带来的洞察能促进人工智能赋能科学发现的进一步发展。

原文摘要 · Abstract (English)

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this discovery process in realistic workflows. However, current scientific benchmarks mostly focus on evaluating the knowledge understanding capabilities of MLLMs, leading to an inadequate assessment of their perception and reasoning abilities. To address this gap, we present the Scientists' First Exam (SFE) benchmark, designed to evaluate the scientific cognitive capacities of MLLMs through three interconnected levels: scientific signal perception, scientific attribute understanding, scientific comparative reasoning. Specifically, SFE comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. Extensive experiments reveal that current state-of-the-art GPT-o3 and InternVL-3 achieve only 34.08% and 26.52% on SFE, highlighting significant room for MLLMs to improve in scientific realms. We hope the insights obtained in SFE will facilitate further developments in AI-enhanced scientific discoveries.

多模态科学推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。