用语言模型做分子逆向设计,让每轮优化都可解释。
Closed-Loop Bayesian Molecular Inverse Design with Semantic LLM Surrogates

- 用冻结的LLM作代理模型,直接理解分子文本和反馈信息。
- 在药物和材料任务中,比单次提示效果更好,接近或超过传统方法。
- 输出可读的决策说明,适合需要透明过程的研究者使用。
实际的分子逆向设计很少是一次性生成的问题,通常采用闭环候选池增强的形式,在有限的实验预算下目标是提高生成分子中符合期望性质比例。贝叶斯优化(BO)为此提供了自然框架,但标准高斯过程代理通常在压缩的连续嵌入空间中运行,丢失了化学家用于判断下一步方向的子结构和参考相似性信号。我们提出 extbf{ extit{method}},一种闭环框架,将代理模型作为设计选择的核心,并用一个冻结的大语言模型实现,该模型直接在原始文本形式的任务指令、SMILES级优化历史和实验反馈上进行推理。每轮迭代中,代理返回结构化决策信号,依据探索与利用原则选择有信息量的参考分子,可选地附带简洁指导语句。该信号转化为冻结分子生成器的下一轮条件文本,生成可检查的自然语言优化轨迹。在MolQA药物与材料设计任务上的实验表明, extit{method}优于单次提示,与基于GP的BO基线相当或更优,并揭示出领域依赖的接口:仅使用参考分子迁移对二元药物靶点效果最佳,而添加简短代理摘要对连续材料任务更有利。
原文摘要 · Abstract (English)
Practical molecular inverse design is rarely a one-shot generation problem; it often takes the form of closed-loop candidate-pool enrichment, where under a limited oracle budget the goal is to \emph{increase the fraction of generated molecules that match a desired property profile}. Bayesian optimization (BO) offers a natural framework for this setting, yet standard Gaussian-process surrogates typically operate in compressed continuous embeddings, which discard the substructural and reference-similarity signals that chemists naturally use to decide where to look next. We propose \textbf{\method}, a closed-loop framework in which the surrogate, rather than the generator, is treated as the locus of design choice, and instantiate it with a frozen large language model that reasons directly over the task instruction, SMILES-level optimization history, and oracle feedback in their native textual form. At each iteration, the surrogate returns a structured decision signal that selects informative reference molecules under an exploration and exploitation principle, optionally with a concise guidance sentence. This signal is converted into next-round conditioning text for a frozen molecular generator, yielding an inspectable optimization trace in natural language. Experiments on MolQA drug and material design tasks show that \method improves over one-shot prompting, is competitive with or stronger than GP-based BO baselines, and reveals a domain-dependent interface: reference-only transfer works best for binary drug targets, while adding a concise surrogate summary is more beneficial for continuous material
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。