arXiv:2503.17656q-bio.QMcs.AI2025-03中稿 · Nature Machine Int…被引 4

为天然产物设计预训练基础模型,提升药物发现效率

Pretraining a Foundation Model for Small-Molecule Natural Products

  • 采用对比学习与掩码图学习结合的预训练策略
  • 在分类、筛选等任务中达当前最优性能
  • 适合天然产物挖掘与新药研发人员使用

天然产物作为微生物、动物或植物的代谢物,具有多样化的生物活性,是药物发现的关键。现有深度学习方法多依赖针对特定下游任务的监督学习,缺乏泛化能力且性能提升空间大。同时,现有分子表征方法不适用于天然产物的独特任务。为此,我们基于天然产物特性预训练了一个基础模型。该方法结合对比学习与掩码图学习目标,强调分子骨架的演化信息并捕捉侧链特征。框架在多种天然产物挖掘与药物发现任务中达到当前最优(SOTA)结果。首先,与合成分子基准相比,其在分类任务中表现更优,证明现有模型难以理解天然合成机制;进一步在基因与微生物层面进行细粒度分析,显示NaFM能有效捕捉演化信息;最后,在虚拟筛选实验中验证了其生成的天然产物表示具有高信息量,可更有效地识别潜在药物候选物。

原文摘要 · Abstract (English)

Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug discovery. Nowadays, existing deep learning methods for natural products research primarily rely on supervised learning approaches designed for specific downstream tasks. However, such one-model-for-a-task paradigm often lacks generalizability and leaves significant room for performance improvement. Additionally, existing molecular characterization methods are not well-suited for the unique tasks associated with natural products. To address these limitations, we have pre-trained a foundation model for natural products based on their unique properties. Our approach employs a novel pretraining strategy that is especially tailored to natural products. By incorporating contrastive learning and masked graph learning objectives, we emphasize evolutional information from molecular scaffolds while capturing side-chain information. Our framework achieves state-of-the-art (SOTA) results in various downstream tasks related to natural product mining and drug discovery. We first compare taxonomy classification with synthesized molecule-focused baselines to demonstrate that current models are inadequate for understanding natural synthesis. Furthermore, by diving into a fine-grained analysis at both the gene and microbial levels, NaFM demonstrates the ability to capture evolutionary information. Eventually, our method is experimented with virtual screening, illustrating informative natural product representations that can lead to more effective identification of potential drug candidates.

天然产物基础模型药物发现分子表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。