arXiv:2505.17873cs.CLcs.AI2025-05被引 6

用模拟实验反馈优化科学假说排序,提升发现效率。

MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback

  • 基于领域知识构建模拟器,通过反馈动态调整假说优先级。
  • 在124个假说上验证,排序趋势与真实实验高度一致。
  • 适合需要高效筛选假设的实验科学领域研究者。

假说排序对自动化科学发现至关重要,尤其在成本高、通量受限的自然科学研究中。现有方法仅依赖语言模型事前推理,缺乏实验反馈。本文提出实验引导排序,根据前期测试反馈优先排序假说。由于真实实验不切实际,我们设计了一个基于领域概念的模拟器,将假说表现建模为与隐藏真实值相似度的函数,并加入噪声扰动。在124个具有实验结果的假说上验证,该模拟器的排序趋势与真实结果保持一致。尽管存在偏差,但其分布类似湿实验噪声,有助于训练更鲁棒的排序策略。我们将实验引导排序建模为序列决策问题,提出一种上下文内强化学习(ICRL)框架。基于大模型的策略将假说分解为功能单元,按机制角色聚类,并依据反馈优先重组。实验表明,该方法显著优于事前基线和强型消融模型。我们的工具包包含模拟器与ICRL框架,支持系统性研究实验引导排序,策略本身亦为有力的可行性证明。

原文摘要 · Abstract (English)

Hypothesis ranking is vital for automated scientific discovery, especially in cost-intensive, throughput-limited natural science domains. Current methods focus on pre-experiment ranking, relying solely on language model reasoning without empirical feedback. We introduce experiment-guided ranking, which prioritizes hypotheses based on feedback from prior tests. Due to the impracticality of real experiments, we propose a simulator grounded in domain-specific concepts that models hypothesis performance as a function of similarity to a hidden ground truth, perturbed by noise. Validated against 124 hypotheses with experimentally reported outcomes, the simulator approximates real results with consistent trend alignment. Although deviations exist, they mimic wet-lab noise, promoting more robust ranking strategies. We frame experiment-guided ranking as a sequential decision-making problem and propose an in-context reinforcement learning (ICRL) framework. Our LLM-based policy decomposes hypotheses into functional elements, clusters them by mechanistic roles, and prioritizes recombinations based on feedback. Experiments show our approach significantly outperforms pre-experiment baselines and strong ablations. Our toolkit, comprising the simulator and ICRL framework, enables systematic research on experiment-guided ranking, with the policy serving as a strong proof of concept.

科学发现强化学习模拟实验假说排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。