测试大模型能否预测科学实验结果,发现其准确率远低于人类专家。
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
- 构建405个跨物理、生物、化学领域的实验预测任务
- 模型准确率14-26%,人类专家约20%,但人类更会判断自身可靠性
- 揭示预测准确性与可信度需同步提升,才可辅助科研决策
加速科学发现需要在投入资源进行昂贵的物理验证前,识别哪些实验可能取得最佳结果。尽管现有基准评估大语言模型(LLMs)在科学知识和推理方面的能力,但它们对实验结果的预测能力——这一人工智能可能显著超越人类的任务——仍缺乏充分探索。我们提出SciPredict,一个包含405个任务的基准,源自33个自然科学子领域的近期实证研究。该基准旨在回答两个关键问题:(a) LLMs能否以足够高的准确率预测科学实验的结果?(b) 这类预测能否可靠地应用于科学研究过程?评估显示,两者均存在根本性局限。模型准确率为14-26%,人类专家表现约为20%。尽管某些前沿模型已超过人类,但整体准确率仍远不足以支撑可靠的实验指导。即使在有限性能下,模型也无法区分可靠与不可靠预测,无论置信度高低或是否认为无需实验即可预测,准确率均仅约20%。相比之下,人类专家表现出强校准能力:随着他们判断结果可预测性增强,准确率从约5%提升至约80%。SciPredict建立了一个严格框架,表明超人类水平的实验科学能力不仅需要更好的预测,还需要对预测可靠性有更准确的认知。为保证可复现性,所有数据与代码均已公开于https://github.com/scaleapi/scipredict。
原文摘要 · Abstract (English)
Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their ability to predict experimental outcomes - a task where AI could significantly exceed human capabilities - remains largely underexplored. We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry. SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process? Evaluations reveal fundamental limitations on both fronts. Model accuracies are 14-26% and human expert performance is $\approx$20%. Although some frontier models exceed human performance model accuracy is still far below what would enable reliable experimental guidance. Even within the limited performance, models fail to distinguish reliable predictions from unreliable ones, achieving only $\approx$20% accuracy regardless of their confidence or whether they judge outcomes as predictable without physical experimentation. Human experts, in contrast, demonstrate strong calibration: their accuracy increases from $\approx$5% to $\approx$80% as they deem outcomes more predictable without conducting the experiment. SciPredict establishes a rigorous framework demonstrating that superhuman performance in experimental science requires not just better predictions, but better awareness of prediction reliability. For reproducibility all our data and code are provided at https://github.com/scaleapi/scipredict
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。