用文本引导扩散模型优化分子多属性,减少预测误差。
Text-guided multi-property molecular optimization with a diffusion language model
- 用文本描述分子属性,通过扩散模型生成目标分子。
- 在基准数据集上同时保持结构相似性和提升化学性质。
- 适合药物研发中需多属性优化的场景。
分子优化(MO)是药物发现中的关键步骤,旨在生成符合工业实际需求的任务导向分子。现有主流方法主要依赖外部性质预测器迭代优化分子属性,但预测器难以覆盖广阔的化学空间,导致预测误差与噪声不可避免,引发误差累积、泛化能力下降和次优分子候选。本文提出一种基于Transformer的扩散语言模型(TransDLM)的文本引导多属性分子优化方法。TransDLM利用标准化化学命名作为分子语义表示,并将性质要求隐式嵌入文本描述中,从而在扩散过程中缓解误差传播。通过融合物理化学细节文本语义与专业分子表示,TransDLM有效整合多种信息源,实现精准优化,增强模型在结构保留与性质提升之间的平衡能力。案例研究进一步验证了其解决实际问题的能力。实验表明,该方法在基准数据集上优于现有最先进方法,在保持分子结构相似性的同时显著提升化学性质。
原文摘要 · Abstract (English)
Molecular optimization (MO) is a crucial stage in drug discovery in which task-oriented generated molecules are optimized to meet practical industrial requirements. Existing mainstream MO approaches primarily utilize external property predictors to guide iterative property optimization. However, learning all molecular samples in the vast chemical space is unrealistic for predictors. As a result, errors and noise are inevitably introduced during property prediction due to the nature of approximation. This leads to discrepancy accumulation, generalization reduction and suboptimal molecular candidates. In this paper, we propose a text-guided multi-property molecular optimization method utilizing transformer-based diffusion language model (TransDLM). TransDLM leverages standardized chemical nomenclature as semantic representations of molecules and implicitly embeds property requirements into textual descriptions, thereby mitigating error propagation during diffusion process. By fusing physically and chemically detailed textual semantics with specialized molecular representations, TransDLM effectively integrates diverse information sources to guide precise optimization, which enhances the model's ability to balance structural retention and property enhancement. Additionally, the success of a case study further demonstrates TransDLM's ability to solve practical problems. Experimentally, our approach surpasses state-of-the-art methods in maintaining molecular structural similarity and enhancing chemical properties on the benchmark dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。