针对药物分子活性悬崖问题,提出半监督学习新方法提升预测准确率。
A Semi-supervised Molecular Learning Framework for Activity Cliff Estimation
- 用未标注数据生成伪标签,结合教师模型评估可信度。
- 在30个数据集上显著提升图神经网络对活性悬崖的预测性能。
- 适合低数据量下分子性质预测,尤其关注活性突变场景。
机器学习可实现快速精准的分子性质预测,广泛应用于药物发现与材料设计。其核心假设是相似分子具有相近性质,但活性悬崖现象会破坏这一假设,导致现有图神经网络等方法性能急剧下降。为应对低数据场景下的挑战,本文提出一种新型半监督学习框架SemiMol,利用大量未标注数据生成伪标签作为后续训练信号。为解决回归任务中无法获取置信度的问题,引入额外教师模型评估伪标签的准确性和可靠性。同时设计自适应课程学习算法,以可控节奏逐步引导目标模型学习困难样本。在30个活性悬崖数据集上的实验表明,SemiMol显著提升了图基模型性能,优于当前主流预训练与半监督基线方法。
原文摘要 · Abstract (English)
Machine learning (ML) enables accurate and fast molecular property predictions, which are of interest in drug discovery and material design. Their success is based on the principle of similarity at its heart, assuming that similar molecules exhibit close properties. However, activity cliffs challenge this principle, and their presence leads to a sharp decline in the performance of existing ML algorithms, particularly graph-based methods. To overcome this obstacle under a low-data scenario, we propose a novel semi-supervised learning (SSL) method dubbed SemiMol, which employs predictions on numerous unannotated data as pseudo-signals for subsequent training. Specifically, we introduce an additional instructor model to evaluate the accuracy and trustworthiness of proxy labels because existing pseudo-labeling approaches require probabilistic outputs to reveal the model's confidence and fail to be applied in regression tasks. Moreover, we design a self-adaptive curriculum learning algorithm to progressively move the target model toward hard samples at a controllable pace. Extensive experiments on 30 activity cliff datasets demonstrate that SemiMol significantly enhances graph-based ML architectures and outpasses state-of-the-art pretraining and SSL baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。