arXiv:2602.13958cs.LGcs.AI2026-02

用100万天然产物训练新型化学语言模型,提升分子生成与性质预测效果。

Chemical Language Models for Natural Products: A State-Space Model Approach

  • 基于状态空间模型Mamba和Transformer对比,针对天然产物设计专用语言模型
  • Mamba生成有效且独特的分子比例高出1-2%,在膜渗透性等任务上性能更优
  • 适用于药物发现中的天然产物生成与性质预测,尤其适合小样本场景

语言模型广泛应用于化学领域,用于分子性质预测与小分子生成,但天然产物(NPs)仍缺乏系统研究,尽管其在药物发现中至关重要。为填补这一空白,我们通过预训练状态空间模型(Mamba和Mamba-2),并与Transformer基线(GPT)进行比较,构建了专用于天然产物的化学语言模型(NPCLMs)。基于约100万天然产物的数据集,首次系统比较了选择性状态空间模型与Transformer在天然产物任务上的表现,并测试了八种分词策略,包括字符级、原子在SMILES中的表示(AIS)、字节对编码(BPE)及专为天然产物设计的BPE。评估指标涵盖分子生成的有效性、唯一性和新颖性,以及膜渗透性、味道和抗癌活性的性质预测,使用MCC和AUC-ROC。结果显示,Mamba比Mamba-2和GPT多生成1-2%的有效且唯一的分子,长程依赖错误更少;而GPT生成的结构略具新颖性。在随机划分下,Mamba变体的MCC优于GPT 0.02-0.04;而在骨架划分下性能相当。结果表明,对约100万天然产物进行领域特化预训练,可达到在超百倍大数据库上训练模型的效果。

原文摘要 · Abstract (English)

Language models are widely used in chemistry for molecular property prediction and small-molecule generation, yet Natural Products (NPs) remain underexplored despite their importance in drug discovery. To address this gap, we develop NP-specific chemical language models (NPCLMs) by pre-training state-space models (Mamba and Mamba-2) and comparing them with transformer baselines (GPT). Using a dataset of about 1M NPs, we present the first systematic comparison of selective state-space models and transformers for NP-focused tasks, together with eight tokenization strategies including character-level, Atom-in-SMILES (AIS), byte-pair encoding (BPE), and NP-specific BPE. We evaluate molecule generation (validity, uniqueness, novelty) and property prediction (membrane permeability, taste, anti-cancer activity) using MCC and AUC-ROC. Mamba generates 1-2 percent more valid and unique molecules than Mamba-2 and GPT, with fewer long-range dependency errors, while GPT yields slightly more novel structures. For property prediction, Mamba variants outperform GPT by 0.02-0.04 MCC under random splits, while scaffold splits show comparable performance. Results demonstrate that domain-specific pre-training on about 1M NPs can match models trained on datasets over 100 times larger.

天然产物语言模型分子生成Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。