arXiv:2409.15370cs.LGcs.AI2024-09被引 6

提出新分词方法,让分子模型覆盖更全、更准确。

Tokenization for Molecular Foundation Models

  • 用n-gram模型快速评估34种分词器,发现现有方法覆盖不足
  • 新提出的Smirk和Smirk-GPE分词器实现OpenSMILES全覆盖
  • 适合做药物、材料等分子设计的科研人员参考

基于文本的基础模型在科学发现中日益重要,分子基础模型推动了材料科学与分子设计的发展。然而,现有模型受限于封闭词汇表的分词器,仅能覆盖分子空间的一小部分。本文系统评估了34种分词器(含19种化学专用分词器),揭示其在SMILES分子表示上的显著覆盖率差距。为评估分词器的影响,引入n-gram语言模型作为低成本代理,并通过预训练和微调18个RoBERTa风格编码器进行分子属性预测验证其有效性。为克服现有分词器局限,提出两种新分词器——Smirk与Smirk-GPE,实现对OpenSMILES规范的完整覆盖。所提方法系统整合原子核、电子与几何自由度,适用于药理学、农业、生物学及能源存储等领域。结果强调了开放词汇建模与化学多样性基准在化学生物信息学中的必要性。

原文摘要 · Abstract (English)

Text-based foundation models have become an important part of scientific discovery, with molecular foundation models accelerating advancements in material science and molecular design.However, existing models are constrained by closed-vocabulary tokenizers that capture only a fraction of molecular space. In this work, we systematically evaluate 34 tokenizers, including 19 chemistry-specific ones, and reveal significant gaps in their coverage of the SMILES molecular representation. To assess the impact of tokenizer choice, we introduce n-gram language models as a low-cost proxy and validate their effectiveness by pretraining and finetuning 18 RoBERTa-style encoders for molecular property prediction. To overcome the limitations of existing tokenizers, we propose two new tokenizers -- Smirk and Smirk-GPE -- with full coverage of the OpenSMILES specification. The proposed tokenizers systematically integrate nuclear, electronic, and geometric degrees of freedom; facilitating applications in pharmacology, agriculture, biology, and energy storage. Our results highlight the need for open-vocabulary modeling and chemically diverse benchmarks in cheminformatics.

分子建模分词器基础模型SMILES

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。