arXiv:2512.13935cs.LGcs.AI2025-12

用大模型先验提升分子发现的贝叶斯优化效率

Informing Acquisition Functions via Foundation Models for Molecular Discovery

  • 不建显式代理模型,直接用大模型先验指导选择分子
  • 通过树状分区和聚类,显著提升大规模候选集的搜索效率
  • 适合数据少、空间大的分子设计场景,尤其适合医药研发

贝叶斯优化(BO)通过构建分子与属性间的概率映射,加速分子发现。传统方法依赖迭代更新代理模型并优化获取函数,但在数据稀疏、候选空间巨大的情况下性能受限。大型语言模型(LLMs)和化学基础模型虽能提供丰富先验,但高维特征、昂贵的上下文学习及深层贝叶斯代理的计算负担制约其应用。为此,我们提出一种无需似然的贝叶斯优化方法,绕过显式建模,直接利用通用LLM和化学基础模型的先验信息指导获取函数。该方法还学习分子搜索空间的树状分区,并在局部使用获取函数,结合蒙特卡洛树搜索实现高效候选选择。进一步引入粗粒度的LLM聚类,将获取函数评估限制在统计上更优的簇内,大幅提高对大规模候选集的可扩展性。大量实验与消融分析表明,该方法在基于大模型引导的分子发现中显著提升了可扩展性、鲁棒性和样本效率。

原文摘要 · Abstract (English)

Bayesian Optimization (BO) is a key methodology for accelerating molecular discovery by estimating the mapping from molecules to their properties while seeking the optimal candidate. Typically, BO iteratively updates a probabilistic surrogate model of this mapping and optimizes acquisition functions derived from the model to guide molecule selection. However, its performance is limited in low-data regimes with insufficient prior knowledge and vast candidate spaces. Large language models (LLMs) and chemistry foundation models offer rich priors to enhance BO, but high-dimensional features, costly in-context learning, and the computational burden of deep Bayesian surrogates hinder their full utilization. To address these challenges, we propose a likelihood-free BO method that bypasses explicit surrogate modeling and directly leverages priors from general LLMs and chemistry-specific foundation models to inform acquisition functions. Our method also learns a tree-structured partition of the molecular search space with local acquisition functions, enabling efficient candidate selection via Monte Carlo Tree Search. By further incorporating coarse-grained LLM-based clustering, it substantially improves scalability to large candidate sets by restricting acquisition function evaluations to clusters with statistically higher property values. We show through extensive experiments and ablations that the proposed method substantially improves scalability, robustness, and sample efficiency in LLM-guided BO for molecular discovery.

分子发现贝叶斯优化大模型生成设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。