arXiv:2607.28304cs.LG2026-07被引 1

用集成共识提升分子图半监督学习效果,无需复杂数据增强。

Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus

论文配图:Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus
图 1 · 摘自论文原文
  • 通过集成模型一致性优化,替代依赖化学敏感增强的常规方法。
  • 在多种分子数据集和GNN架构上均显著提升预测精度,优于传统全监督集成。
  • 提升模型鲁棒性并降低校准误差,适合缺乏标签的分子设计场景。

机器学习正加速分子科学中的性质预测、模拟及新分子与材料发现。然而,获取标注数据成本高且耗时,而大量未标注分子数据却易得。传统半监督学习依赖保持标签不变的数据增强,但在分子领域难以设计,因微小变化可能大幅改变性质。本文表明,基于集成共识的半监督方法可在多种分子数据集、任务类型和图神经网络架构下提升预测准确率。训练中采用集成共识目标可增强模型鲁棒性,其效果类似知识蒸馏;单个集成成员经此训练后,在几乎所有情况下均优于传统监督方式训练的完整集成。此外,此类半监督训练还能降低校准误差。

原文摘要 · Abstract (English)

Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.

分子图半监督学习集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。