用真实科研场景评估大模型,发现其科学探索能力仍有显著短板。
Evaluating Large Language Models in Scientific Discovery
- 构建跨学科科研任务框架,模拟真实研究流程
- 模型在复杂项目中表现远低于通用基准,且规模越大提升越小
- 适合关注科研自动化与大模型真实能力的学者
大语言模型(LLMs)在科学研究中的应用日益广泛,但现有科学评测多基于脱离上下文的知识问答,忽视了科学发现中的迭代推理、假设生成与观测解读等核心环节。本文提出一种情景驱动的评测框架(SDE),涵盖生物、化学、材料与物理领域,由领域专家定义真实科研项目,并将其分解为可验证的研究情景,从中采样标准化问题。该框架从两个层面评估:(i) 针对情景相关问题的答题准确率;(ii) 项目级表现,要求模型提出可检验假设、设计仿真或实验并解释结果。对前沿LLMs的测试显示,相比通用科学基准,模型在科学发现任务中存在持续性能差距,模型规模与推理能力的提升收益递减,且不同厂商顶级模型均表现出系统性弱点。研究情景间表现差异显著,导致最佳模型选择随任务变化,表明当前所有模型距离通用科学“超智能”仍相去甚远。然而,即便部分情景得分较低,模型在多种科研项目中仍展现出潜力,凸显引导式探索与偶然发现的重要性。SDE框架为科学发现相关评估提供了可复现的标准,也为推动模型向科学探索演进指明了实践路径。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。