用大模型自动发现化学新假说,效果接近人类专家。
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
- 将假说生成拆解为找灵感、组合假说、排序三步,构建可执行框架。
- 在51篇高影响力论文上复现假说,相似度高且无数据泄露。
- 大模型意外精准捕捉未见科学关联,暗示其可能已隐含人类未知知识。
科学发现推动社会发展,大语言模型(LLMs)或能加速这一进程。然而,现有研究尚不清楚LLMs能否在仅提供研究背景(如问题或综述)的前提下,自主生成新颖且有效的化学假说。本文提出一个数学分解框架:多数化学假说可由研究背景与一组灵感组合而成,据此形成三个可操作子任务——灵感检索、假说生成、假说排序。基于此,构建了代理式框架MOOSE-Chem。为评估该框架,构建了一个包含51篇2024年后发表的高影响力化学论文的基准集,每篇均由博士级化学家标注背景、灵感和假说。结果表明,该框架能以高相似度复现真实假说,准确捕捉核心创新,且因使用知识截止于2024年前的LLM,避免数据污染。此外,模型在具有明显分布外特征的灵感检索任务中表现优异,提示其可能已编码尚未被人类识别的潜在科学知识关联。
原文摘要 · Abstract (English)
Scientific discovery plays a pivotal role in advancing human society, and recent progress in large language models (LLMs) suggests their potential to accelerate this process. However, it remains unclear whether LLMs can autonomously generate novel and valid hypotheses in chemistry. In this work, we investigate whether LLMs can discover high-quality chemistry hypotheses given only a research background-comprising a question and/or a survey-without restriction on the domain of the question. We begin with the observation that hypothesis discovery is a seemingly intractable task. To address this, we propose a formal mathematical decomposition grounded in a fundamental assumption: that most chemistry hypotheses can be composed from a research background and a set of inspirations. This decomposition leads to three practical subtasks-retrieving inspirations, composing hypotheses with inspirations, and ranking hypotheses - which together constitute a sufficient set of subtasks for the overall scientific discovery task. We further develop an agentic LLM framework, MOOSE-Chem, that is a direct implementation of this mathematical decomposition. To evaluate this framework, we construct a benchmark of 51 high-impact chemistry papers published and online after January 2024, each manually annotated by PhD chemists with background, inspirations, and hypothesis. The framework is able to rediscover many hypotheses with high similarity to the groundtruth, successfully capturing the core innovations-while ensuring no data contamination since it uses an LLM with knowledge cutoff date prior to 2024. Finally, based on LLM's surprisingly high accuracy on inspiration retrieval, a task with inherently out-of-distribution nature, we propose a bold assumption: that LLMs may already encode latent scientific knowledge associations not yet recognized by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。