arXiv:2412.15589cs.LGcs.AI2024-12AAAI被引 23

无需人工标注,自动学习分子功能团并生成强区分性表示

Pre-training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information Bottleneck

  • 用子图条件信息瓶颈机制自动发现分子中的关键子结构
  • 在多个分子数据集上实现优于现有方法的预训练性能
  • 适合药物发现、分子性质预测等需自监督学习的场景

本研究旨在构建无需人工标注或先验知识的分子图神经网络预训练模型。尽管已有多种方法尝试克服标注分子获取难的问题,但以往预训练方法仍依赖于语义子图(如官能团),仅关注官能团可能忽略图级别的差异。构建分子预训练GNN的关键挑战在于:(1) 生成具有强区分性的图级表示;(2) 在无先验知识下自动发现官能团。为此,我们提出一种新型子图条件图信息瓶颈方法(S-CGIB),用于识别核心子图(图核心)和显著子图。其核心思想是:图核心包含压缩且充分的信息,可在给定显著子图条件下重构输入图,并生成强区分性的图级表示。为在无先验知识下发现显著子图,我们生成一组官能团候选(即自中心网络),并通过图核心与候选间的注意力交互进行筛选。尽管这些子图通过自监督学习获得,其结果却与真实世界官能团高度吻合。在多个领域分子数据集上的大量实验表明,S-CGIB具有显著优势。

原文摘要 · Abstract (English)

This study aims to build a pre-trained Graph Neural Network (GNN) model on molecules without human annotations or prior knowledge. Although various attempts have been proposed to overcome limitations in acquiring labeled molecules, the previous pre-training methods still rely on semantic subgraphs, i.e., functional groups. Only focusing on the functional groups could overlook the graph-level distinctions. The key challenge to build a pre-trained GNN on molecules is how to (1) generate well-distinguished graph-level representations and (2) automatically discover the functional groups without prior knowledge. To solve it, we propose a novel Subgraph-conditioned Graph Information Bottleneck, named S-CGIB, for pre-training GNNs to recognize core subgraphs (graph cores) and significant subgraphs. The main idea is that the graph cores contain compressed and sufficient information that could generate well-distinguished graph-level representations and reconstruct the input graph conditioned on significant subgraphs across molecules under the S-CGIB principle. To discover significant subgraphs without prior knowledge about functional groups, we propose generating a set of functional group candidates, i.e., ego networks, and using an attention-based interaction between the graph core and the candidates. Despite being identified from self-supervised learning, our learned subgraphs match the real-world functional groups. Extensive experiments on molecule datasets across various domains demonstrate the superiority of S-CGIB.

图神经网络分子建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。