arXiv:2502.19794cs.CEcs.LG2025-02被引 3

用大模型生成满足多种约束的分子,性能超越现有方法。

ChatMol: A Versatile Molecule Designer Based on the Numerically Enhanced Large Language Model

  • 基于大模型设计分子,通过提示工程与反馈学习实现多约束生成
  • 在多目标结合亲和力任务中,KD值低至0.25,性能提升4.76%
  • 引入数值编码增强,属性预测相关性提升0.49,适合药物研发人员

面向药物发现中的目标导向新分子设计(即生成具有特定性质或结构约束的分子),现有方法如贝叶斯优化和强化学习需训练多个性质预测器,且难以融入结构约束。受大语言模型(LLM)在文本生成中的成功启发,我们提出ChatMol,一种利用LLM实现多样化约束条件下分子设计的新方法。首先,构建适配LLM的分子表示并验证其在多个在线LLM上的有效性;其次,设计针对不同约束任务的提示策略,进一步微调现有LLM,并融合来自性质预测的反馈学习;最后,为克服LLM在数值识别上的局限,借鉴位置编码思想,在提示中加入数值额外编码。在单性质、结构-性质及多性质约束任务上的实验表明,ChatMol持续优于当前最优基线(包括VAE和强化学习方法)。特别地,在多目标结合亲和力最大化任务中,对蛋白靶点ESR1,ChatMol实现0.25的显著降低的KD值,整体性能领先前序方法4.76%;同时,经数值增强后,指令属性值与生成分子实际属性间的皮尔逊相关系数最高提升0.49。这些结果凸显了大模型作为通用分子生成框架的巨大潜力,为传统潜在空间与强化学习方法提供了有前景的替代方案。

原文摘要 · Abstract (English)

Goal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints, is a crucial yet challenging task in drug discovery. Existing methods, such as Bayesian optimization and reinforcement learning, often require training multiple property predictors and struggle to incorporate substructure constraints. Inspired by the success of Large Language Models (LLMs) in text generation, we propose ChatMol, a novel approach that leverages LLMs for molecule design across diverse constraint settings. Initially, we crafted a molecule representation compatible with LLMs and validated its efficacy across multiple online LLMs. Afterwards, we developed specific prompts geared towards diverse constrained molecule generation tasks to further fine-tune current LLMs while integrating feedback learning derived from property prediction. Finally, to address the limitations of LLMs in numerical recognition, we referred to the position encoding method and incorporated additional encoding for numerical values within the prompt. Experimental results across single-property, substructure-property, and multi-property constrained tasks demonstrate that ChatMol consistently outperforms state-of-the-art baselines, including VAE and RL-based methods. Notably, in multi-objective binding affinity maximization task, ChatMol achieves a significantly lower KD value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76%. Meanwhile, with numerical enhancement, the Pearson correlation coefficient between the instructed property values and those of the generated molecules increased by up to 0.49. These findings highlight the potential of LLMs as a versatile framework for molecule generation, offering a promising alternative to traditional latent space and RL-based approaches.

分子生成大模型药物设计提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。