arXiv:2508.13408cs.LG2025-08被引 4

基于15亿分子预训练,提升从头生成分子的效率与效果

NovoMolGen: Rethinking Molecular Language Model Pretraining

  • 采用1.5亿分子数据集,基于字符串表示构建Transformer模型
  • 实证发现预训练指标与下游性能相关性弱,揭示分子建模特殊性
  • 在无约束和目标导向生成任务中均达到新基准,适合药物设计应用

设计具有特定性质的新分子需高效探索从10^23到10^60种可能合成分子的巨大化学空间。尽管已有多种深度生成模型用于小分子设计,基于字符串表示的分子大语言模型(Mol-LLMs)因其可扩展性,能够探索数十亿分子。然而,标准语言建模实践如文本表示、分词策略、模型规模和数据集规模对分子生成性能的影响仍不明确。本文通过引入NovoMolGen——一个在15亿分子上预训练的基于Transformer的基座模型,系统研究了这些关键因素。通过广泛的实证分析,我们发现预训练阶段的性能度量与实际下游性能之间存在弱相关性,揭示了分子与通用NLP训练动态的重要差异。NovoMolGen在无约束和目标导向分子生成任务中均显著优于先前的Mol-LLMs及专用生成模型,为高效、有效的分子建模策略提供了坚实基础。

原文摘要 · Abstract (English)

Designing de-novo molecules with desired property profiles requires efficient exploration of the vast chemical space ranging from $10^{23}$ to $10^{60}$ possible synthesizable candidates. While various deep generative models have been developed to design small molecules using diverse input representations, Molecular Large Language Models (Mol-LLMs) based on string representations have emerged as a scalable approach capable of exploring billions of molecules. However, there remains limited understanding regarding how standard language modeling practices such as textual representations, tokenization strategies, model size, and dataset scale impact molecular generation performance. In this work, we systematically investigate these critical aspects by introducing NovoMolGen, a family of transformer-based foundation models pretrained on 1.5 billion molecules for de-novo molecule generation. Through extensive empirical analyses, we identify a weak correlation between performance metrics measured during pretraining and actual downstream performance, revealing important distinctions between molecular and general NLP training dynamics. NovoMolGen establishes new state-of-the-art results, substantially outperforming prior Mol-LLMs and specialized generative models in both unconstrained and goal-directed molecular generation tasks, thus providing a robust foundation for advancing efficient and effective molecular modeling strategies.

分子生成大语言模型药物设计预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。