从分子属性数据中逆向推导药物优化逻辑,让AI学会像专家一样思考。
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
- 通过聚类共享结构片段的分子,用大模型分析结构变化与性质差异的关系。
- 在230万条数据上实现18项任务中15项最优,包括多属性联合优化。
- 可解释推理过程,支持新靶点、新性质等未知场景,适合药物研发人员使用。
新兴的推理模型有望自动化科学发现,但其训练面临关键瓶颈:实验结果丰富,而中间推理步骤却极少被大规模记录。为此,我们提出DESRO框架,从结果数据中解码科学推理过程。通过分析分组数据中的共性模式与关键差异,大型语言模型(LLM)可还原背后的逻辑。我们在药物分子优化这一关键环节中实现该框架,传统上依赖药化专家的迭代推理。基于230万条分子性质记录,框架通过聚类共享片段的分子,再利用LLM分析结构变异与性质变化的相关性,进而推导出优化逻辑。据此训练的模型能以可解释的方式进行分子优化。DESRO在18项任务中取得15项最高成功率,涵盖生物活性与ADMET性质的单/多属性优化。其推理过程具备强泛化能力,适用于分布外场景,包括新型性质组合、未见过的生物靶点及仅由自然语言定义的新性质。在严格时间划分的回顾性案例研究中,模型能自主重构专家级先导化合物优化路径。此外,该框架还可拓展至反应配体选择。结果表明,从结果数据中解码推理步骤是一种可行且可扩展的科学推理范式,为加速科学发现提供新路径。
原文摘要 · Abstract (English)
Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are rarely documented at scale. To bridge this gap, we propose DESRO, a framework for deciphering scientific reasoning from outcomes. By analyzing shared patterns and key differences within grouped data, a large language model (LLM) can recover the underlying logic. We instantiate this framework in molecule optimization, a pivotal stage in drug discovery that traditionally relies on the iterative reasoning of medicinal chemists. Across 2.3 million molecular property records, our framework infers optimization rationales by grouping molecules with shared fragments, then using an LLM to analyze how structural variations correlate with property differences. Based on the derived data, we train a model that conducts molecule optimization through an interpretable reasoning process. DESRO achieves the highest success rates on 15 out of 18 tasks, spanning both single- and multi-property optimization of bioactivity and ADMET properties. The reasoning process enables robust generalization to out-of-distribution scenarios, including novel property combinations, unseen biological targets, and unseen properties defined solely by natural language descriptions. In retrospective case studies under strict temporal splits, the model autonomously reconstructs expert-level lead optimization trajectories. Additionally, our framework extends beyond molecule optimization to reaction ligand selection. Our results establish deciphering reasoning steps from outcome data as a viable paradigm for enabling scientific reasoning, providing a scalable approach to accelerate scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。