arXiv:2602.20344cs.LGcs.AI2026-02被引 1

通过碎片化自监督学习,让模型更懂分子结构的化学意义。

Hierarchical Molecular Representation Learning via Fragment-Based Self-Supervised Embedding Prediction

  • 将分子拆解为化学有意义的碎片,分层学习原子与碎片表示
  • 在多个数据集上超越现有自监督方法,尤其在迁移学习中表现优异
  • 适合需要理解分子结构的药物发现与材料设计场景

图自监督学习(GSSL)在无需人工标注的情况下生成具有表达力的图嵌入,对标签成本高的分子图分析尤为有价值。然而,现有GSSL方法多关注节点或边级信息,常忽略影响分子性质的化学子结构。本文提出图语义预测网络(GraSPNet),一种分层自监督框架,显式建模原子级与碎片级语义。GraSPNet无需预定义词汇表即可将分子图分解为化学相关的碎片,并通过双层级消息传递与掩码语义预测学习节点和碎片级表示。该分层语义监督使GraSPNet能捕捉多层次结构信息,兼具表达力与可迁移性。在多个分子属性预测基准上的实验表明,GraSPNet学习到的表示具有化学意义,在迁移学习设置下持续优于当前最优的GSSL方法。

原文摘要 · Abstract (English)

Graph self-supervised learning (GSSL) has demonstrated strong potential for generating expressive graph embeddings without the need for human annotations, making it particularly valuable in domains with high labeling costs such as molecular graph analysis. However, existing GSSL methods mostly focus on node- or edge-level information, often ignoring chemically relevant substructures which strongly influence molecular properties. In this work, we propose Graph Semantic Predictive Network (GraSPNet), a hierarchical self-supervised framework that explicitly models both atomic-level and fragment-level semantics. GraSPNet decomposes molecular graphs into chemically meaningful fragments without predefined vocabularies and learns node- and fragment-level representations through multi-level message passing with masked semantic prediction at both levels. This hierarchical semantic supervision enables GraSPNet to learn multi-resolution structural information that is both expressive and transferable. Extensive experiments on multiple molecular property prediction benchmarks demonstrate that GraSPNet learns chemically meaningful representations and consistently outperforms state-of-the-art GSSL methods in transfer learning settings.

分子表示自监督学习图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。