arXiv:2411.03537cs.LGcs.AI2024-11被引 2

用两阶段预训练提升真实世界分子属性预测能力

Two-Stage Pretraining for Molecular Property Prediction in the Wild

  • 先通过掩码原子和极端去噪学习无标签分子表示
  • 再用计算方法生成的辅助属性优化表示,跨22个数据集表现领先
  • 适合标签稀缺场景,尤其对实验成本高的分子研究者有用

分子深度学习模型在属性预测上已取得显著进展,但通常需要大量标注数据。然而在真实应用中,标签极度稀缺,因实验验证成本高、耗时长。本文提出MoleVers,一种适用于多种分子属性预测的通用预训练模型,针对标签稀缺的现实场景。MoleVers采用两阶段预训练:第一阶段利用新提出的分支编码器架构与动态噪声尺度采样,通过掩码原子预测和极端去噪任务从无标签数据中学习分子表征;第二阶段通过密度泛函理论或大语言模型生成的辅助属性进一步优化表征。在22个小型实验验证数据集上的评估表明,MoleVers达到当前最佳性能,验证了其两阶段框架在生成多样化下游任务可泛化表征方面的有效性。

原文摘要 · Abstract (English)

Molecular deep learning models have achieved remarkable success in property prediction, but they often require large amounts of labeled data. The challenge is that, in real-world applications, labels are extremely scarce, as obtaining them through laboratory experimentation is both expensive and time-consuming. In this work, we introduce MoleVers, a versatile pretrained molecular model designed for various types of molecular property prediction in the wild, i.e., where experimentally-validated labels are scarce. MoleVers employs a two-stage pretraining strategy. In the first stage, it learns molecular representations from unlabeled data through masked atom prediction and extreme denoising, a novel task enabled by our newly introduced branching encoder architecture and dynamic noise scale sampling. In the second stage, the model refines these representations through predictions of auxiliary properties derived from computational methods, such as the density functional theory or large language models. Evaluation on 22 small, experimentally-validated datasets demonstrates that MoleVers achieves state-of-the-art performance, highlighting the effectiveness of its two-stage framework in producing generalizable molecular representations for diverse downstream properties.

分子建模预训练小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。