评测大模型在真实科研中的推理能力,发现其仍难可靠推断实验结果。
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

- 构建多模态科研评测基准SEE,基于真实文献与实验设计
- 顶尖模型准确率仅48.7%,工具使用提升至52.7%但未解决根本问题
- 强调科学推理需基于证据边界,适合关注AI科研辅助的研究者
大型语言模型(LLMs)在科学发现中应用日益广泛,但其能否支持复杂真实的实验室科学研究仍不明确。本文提出科学边缘评估(Science Edge Evaluation, SEE),一个基于同行评审文献和化学、生物、材料科学实验实践的专家标注多模态基准。对19个多模态大语言模型(MLLMs)的评估显示,即使表现最佳的模型准确率也仅为48.7%。此外,通用模型平均优于专用科学模型。在视觉代理评估中,使用工具将最优准确率提升至52.7%。工具可扩展模型信息获取能力,但更多信息并不必然带来可靠的科学推理。关键挑战在于模型能否在原始实验证据范围内管理工具生成的信息。这些发现表明,当前MLLMs仍无法从实验结果中做出合理且受证据约束的推断,而这正是真实科学发现的核心能力。弥合这一差距需要MLLMs从解释已有科学概念转向从实验数据中推导新颖且有证据支持的见解。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。