用100万天然产物训练新型化学语言模型,提升分子生成与性质预测效果。
Chemical Language Models for Natural Products: A State-Space Model Approach
- 基于状态空间模型Mamba和Transformer对比,针对天然产物设计专用语言模型
- Mamba生成有效且独特的分子比例高出1-2%,在膜渗透性等任务上性能更优
- 适用于药物发现中的天然产物生成与性质预测,尤其适合小样本场景
语言模型广泛应用于化学领域,用于分子性质预测与小分子生成,但天然产物(NPs)仍缺乏系统研究,尽管其在药物发现中至关重要。为填补这一空白,我们通过预训练状态空间模型(Mamba和Mamba-2),并与Transformer基线(GPT)进行比较,构建了专用于天然产物的化学语言模型(NPCLMs)。基于约100万天然产物的数据集,首次系统比较了选择性状态空间模型与Transformer在天然产物任务上的表现,并测试了八种分词策略,包括字符级、原子在SMILES中的表示(AIS)、字节对编码(BPE)及专为天然产物设计的BPE。评估指标涵盖分子生成的有效性、唯一性和新颖性,以及膜渗透性、味道和抗癌活性的性质预测,使用MCC和AUC-ROC。结果显示,Mamba比Mamba-2和GPT多生成1-2%的有效且唯一的分子,长程依赖错误更少;而GPT生成的结构略具新颖性。在随机划分下,Mamba变体的MCC优于GPT 0.02-0.04;而在骨架划分下性能相当。结果表明,对约100万天然产物进行领域特化预训练,可达到在超百倍大数据库上训练模型的效果。
原文摘要 · Abstract (English)
Language models are widely used in chemistry for molecular property prediction and small-molecule generation, yet Natural Products (NPs) remain underexplored despite their importance in drug discovery. To address this gap, we develop NP-specific chemical language models (NPCLMs) by pre-training state-space models (Mamba and Mamba-2) and comparing them with transformer baselines (GPT). Using a dataset of about 1M NPs, we present the first systematic comparison of selective state-space models and transformers for NP-focused tasks, together with eight tokenization strategies including character-level, Atom-in-SMILES (AIS), byte-pair encoding (BPE), and NP-specific BPE. We evaluate molecule generation (validity, uniqueness, novelty) and property prediction (membrane permeability, taste, anti-cancer activity) using MCC and AUC-ROC. Mamba generates 1-2 percent more valid and unique molecules than Mamba-2 and GPT, with fewer long-range dependency errors, while GPT yields slightly more novel structures. For property prediction, Mamba variants outperform GPT by 0.02-0.04 MCC under random splits, while scaffold splits show comparable performance. Results demonstrate that domain-specific pre-training on about 1M NPs can match models trained on datasets over 100 times larger.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。