arXiv:2507.10273cs.LGq-bio.BM2025-07被引 2

用生物背景条件生成药物分子,让设计更精准高效。

Conditional Chemical Language Models are Versatile Tools in Drug Discovery

  • 基于片段序列建模,根据靶点等生物信息生成分子
  • 零样本测试中性能优于或持平现有方法,速度更快
  • 可解释性强,适合早期药物研发人员使用

生成式化学语言模型在分子设计中表现强劲,但其在药物发现中的应用受限于缺乏可靠的奖励信号和输出不可解释。我们提出SAFE-T,一个通用化学建模框架,通过蛋白质靶点或作用机制等生物上下文条件,无需依赖结构信息或人工评分函数,即可优先筛选和设计分子。SAFE-T对给定生物提示下片段化分子序列的条件概率进行建模,实现虚拟筛选、药物-靶点相互作用预测和活性悬崖检测等任务的合理打分。同时支持目标导向生成,采样自学习分布以对齐生物学目标。在包含预测(LIT-PCBA、DAVIS、KIBA、ACNet)和生成(DRUG、PMO)基准的全面零样本评估中,SAFE-T性能持续优于或相当现有方法,且显著更快。片段级归因分析表明,SAFE-T捕捉到已知的构效关系,支持可解释且生物学合理的分子设计。结合计算效率,结果表明条件生成式化学语言模型可统一打分与生成,加速早期药物发现。

原文摘要 · Abstract (English)

Generative chemical language models (CLMs) have demonstrated strong capabilities in molecular design, yet their impact in drug discovery remains limited by the absence of reliable reward signals and the lack of interpretability in their outputs. We present SAFE-T, a generalist chemical modeling framework that conditions on biological context -- such as protein targets or mechanisms of action -- to prioritize and design molecules without relying on structural information or engineered scoring functions. SAFE-T models the conditional likelihood of fragment-based molecular sequences given a biological prompt, enabling principled scoring of molecules across tasks such as virtual screening, drug-target interaction prediction, and activity cliff detection. Moreover, it supports goal-directed generation by sampling from this learned distribution, aligning molecular design with biological objectives. In comprehensive zero-shot evaluations across predictive (LIT-PCBA, DAVIS, KIBA, ACNet) and generative (DRUG, PMO) benchmarks, SAFE-T consistently achieves performance comparable to or better than existing approaches while being significantly faster. Fragment-level attribution further reveals that SAFE-T captures known structure-activity relationships, supporting interpretable and biologically grounded design. Together with its computational efficiency, these results demonstrate that conditional generative CLMs can unify scoring and generation to accelerate early-stage drug discovery.

药物发现生成模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。