arXiv:2411.01401q-bio.BMcs.AI2024-11被引 6

通过随机遮蔽分子式中的功能基团,提升语言模型对分子结构的理解能力。

Pre-trained Molecular Language Models with Random Functional Group Masking

  • 在预训练时随机遮蔽SMILES中对应的功能基团序列
  • 在11项任务中9项超越现有模型,表现最优
  • 适合药物发现与材料科学中的分子性质预测

计算化学领域近年来利用基于Transformer的分子语言模型(如MoLFormer),通过大量简化分子输入线性系统(SMILES)序列进行预训练,以理解并预测分子属性与活性,这对药物研发和材料科学至关重要。为提升性能,研究者引入图神经网络(如GEM),将分子拓扑、几何及二维或三维结构纳入预训练过程。尽管现有方法大多自动将SMILES转换为分子图,但本研究假设基于SMILES的模型可能隐式学习到结构感知表示。本文提出 extit{ours}——一种基于SMILES的分子语言模型,在预训练阶段随机遮蔽对应特定功能基团的子序列,以融入原子结构信息。该方法旨在促使模型更准确推断分子结构与性质,从而增强预测能力。在化学领域的11个基准分类与回归任务上进行的广泛实验表明, extit{ours}具有鲁棒性和优越性,其在11项任务中9项超越现有基于SMILES或图的预训练模型,其余两项排名紧随第二。

原文摘要 · Abstract (English)

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to understand and predict molecular properties and activities, a critical step in fields like drug discovery and materials science. To further improve performance, researchers have introduced graph neural networks with graph-based molecular representations, such as GEM, incorporating the topology, geometry, 2D or even 3D structures of molecules into pre-training. While most of molecular graphs in existing studies were automatically converted from SMILES sequences, it is to assume that transformer-based language models might be able to implicitly learn structure-aware representations from SMILES sequences. In this paper, we propose \ours{} -- a SMILES-based \underline{\em M}olecular \underline{\em L}anguage \underline{\em M}odel, which randomly masking SMILES subsequences corresponding to specific molecular \underline{\em F}unctional \underline{\em G}roups to incorporate structure information of atoms during the pre-training phase. This technique aims to compel the model to better infer molecular structures and properties, thus enhancing its predictive capabilities. Extensive experimental evaluations across 11 benchmark classification and regression tasks in the chemical domain demonstrate the robustness and superiority of \ours{}. Our findings reveal that \ours{} outperforms existing pre-training models, either based on SMILES or graphs, in 9 out of the 11 downstream tasks, ranking as a close second in the remaining ones.

分子建模语言模型预训练药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。