用分层搜索让大模型生成可实验的精细科学假说
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
- 通过分层迭代添加细节,逐步构建具体实验方案
- 在专家标注数据集上优于现有基线,显著提升假说质量
- 适合需要自动化科研设计的科学家和AI辅助研究者
大型语言模型(LLMs)在自动提出科学假说方面展现出潜力,但现有方法多生成粗粒度假说,缺乏关键的方法与实验细节。本文首次正式定义细粒度科学假说发现任务,即从粗略研究方向生成可实验执行的详细假说。将该任务建模为组合优化问题,探究在最大利用下大模型的上限能力。我们研究四个核心问题:(1)如何利用模型自身内部启发式,基于其内部评分选出最可能成立的假说,从而定义隐式奖励景观;(2)模型判断更优的假说是否与真实假说更具一致性;(3)使用同容量多模型集合塑造奖励景观是否优于单一最强模型重复使用;(4)相同模型的集合能否提供比单个模型更稳定的奖励景观。为此,我们提出分层搜索方法,逐步从泛化概念推进至具体实验配置。实验证明,该过程使奖励景观更平滑,实现更有效的优化。在新构建的由近期文献专家标注的细粒度假说基准上,本方法持续超越强基线。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothesis discovery, which entails generating detailed, experimentally actionable hypotheses from coarse initial research directions. We frame this as a combinatorial optimization problem and investigate the upper limits of LLMs' capacity to solve it when maximally leveraged. Specifically, we explore four foundational questions: (1) how to best harness an LLM's internal heuristics to formulate the fine-grained hypothesis it itself would judge as the most promising among all the possible hypotheses it might generate, based on its own internal scoring-thus defining a latent reward landscape over the hypothesis space; (2) whether such LLM-judged better hypotheses exhibit stronger alignment with ground-truth hypotheses; (3) whether shaping the reward landscape using an ensemble of diverse LLMs of similar capacity yields better outcomes than defining it with repeated instances of the strongest LLM among them; and (4) whether an ensemble of identical LLMs provides a more reliable reward landscape than a single LLM. To address these questions, we propose a hierarchical search method that incrementally proposes and integrates details into the hypothesis, progressing from general concepts to specific experimental configurations. We show that this hierarchical process smooths the reward landscape and enables more effective optimization. Empirical evaluations on a new benchmark of expert-annotated fine-grained hypotheses from recent literature show that our method consistently outperforms strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。