arXiv:2504.08874cs.LGcs.AI2025-04被引 1

用大模型提取化学知识,加速反应优化

Distilling and exploiting quantitative insights from Large Language Models for enhanced Bayesian optimization of chemical reactions

  • 通过提示工程和偏好学习从大模型中挖掘化学先验知识
  • 零样本下先验函数与实验产率有弱相关性,但能引导优化方向
  • 在6个数据集中的4个提升初始查询表现,加速收敛

机器学习与贝叶斯优化(BO)算法可显著加速化学反应的优化。迁移学习能在数据稀缺情况下,通过利用任务外的已有化学信息(源数据)提升BO效果。大语言模型(LLMs)在基础训练数据中蕴含丰富的化学知识,具备处理化学数据的能力,并可融合多种模态的源数据。本文研究如何从LLMs中提取化学信息,用于迁移学习以加速反应条件的贝叶斯优化。我们发现,采用类似问卷的提示策略与偏好学习,可推断出一个建模化学参数空间先验知识的效用函数;尽管处于零样本设置,该函数与真实实验产率在参数空间上表现出弱相关性。更重要的是,该效用函数能引导BO聚焦于有潜力的区域,提升初始查询的产率,并在6个数据集中的4个实现优化性能提升。本工作为弥合大模型中隐含的化学知识与严谨贝叶斯优化方法之间的差距迈出一步。

原文摘要 · Abstract (English)

Machine learning and Bayesian optimization (BO) algorithms can significantly accelerate the optimization of chemical reactions. Transfer learning can bolster the effectiveness of BO algorithms in low-data regimes by leveraging pre-existing chemical information or data outside the direct optimization task (i.e., source data). Large language models (LLMs) have demonstrated that chemical information present in foundation training data can give them utility for processing chemical data. Furthermore, they can be augmented with and help synthesize potentially multiple modalities of source chemical data germane to the optimization task. In this work, we examine how chemical information from LLMs can be elicited and used for transfer learning to accelerate the BO of reaction conditions to maximize yield. Specifically, we show that a survey-like prompting scheme and preference learning can be used to infer a utility function which models prior chemical information embedded in LLMs over a chemical parameter space; we find that the utility function shows modest correlation to true experimental measurements (yield) over the parameter space despite operating in a zero-shot setting. Furthermore, we show that the utility function can be leveraged to focus BO efforts in promising regions of the parameter space, improving the yield of the initial BO query and enhancing optimization in 4 of the 6 datasets studied. Overall, we view this work as a step towards bridging the gap between the chemistry knowledge embedded in LLMs and the capabilities of principled BO methods to accelerate reaction optimization.

贝叶斯优化大模型应用化学反应迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。