arXiv:2507.22514cs.LG2025-07

用专属分子语言预训练,提升药物属性预测精度与效率

SmilesT5: Domain-specific pretraining for molecular language models

  • 设计分子语言专用的文本到文本预训练任务
  • 在6个分类任务上优于传统训练和先前微调方法
  • 固定嵌入可替代微调,大幅降低计算开销

分子性质预测在药物发现与开发中日益关键。通常神经网络可通过图结构、语言或特征方法学习分子性质。自然语言处理的进展表明,基于掩码语言建模的大规模变换器模型能有效学习复杂人类语言。此类方法也被用于学习以简化分子输入线性系统(SMILES)字符串表示的分子语言。本文提出新的领域专用文本到文本预训练任务,在六个基于分类的分子性质预测基准上表现优于传统似然训练和此前提出的微调任务。消融实验显示,这些预训练任务可提升数据与计算效率。最后,模型预训练嵌入可作为下游机器学习分类器的固定输入,性能媲美微调但计算开销显著更低。

原文摘要 · Abstract (English)

Molecular property prediction is an increasingly critical task within drug discovery and development. Typically, neural networks can learn molecular properties using graph-based, language-based or feature-based methods. Recent advances in natural language processing have highlighted the capabilities of neural networks to learn complex human language using masked language modelling. These approaches to training large transformer-based deep learning models have also been used to learn the language of molecules, as represented by simplified molecular-input line-entry system (SMILES) strings. Here, we present novel domain-specific text-to-text pretraining tasks that yield improved performance in six classification-based molecular property prediction benchmarks, relative to both traditional likelihood-based training and previously proposed fine-tuning tasks. Through ablation studies, we show that data and computational efficiency can be improved by using these domain-specific pretraining tasks. Finally, the pretrained embeddings from the model can be used as fixed inputs into a downstream machine learning classifier and yield comparable performance to finetuning but with much lower computational overhead.

分子语言预训练药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。