构建分层分子图,同时学习原子、键和片段的语义,提升属性预测效果。
Multi-level Self-supervised Pretraining on Compositional Hierarchical Graph for Molecular Property Prediction

- 用四类节点构建三层分层图,原子与键并行建模,独立演化。
- 在九个MoleculeNet数据集上7个领先,整体性能优于现有方法。
- 适合需要精准分子结构理解的药物发现与化学工程研究者。
基于分子图的自监督预训练已成为分子属性预测的有前景方法,但现有多数方法仅在单一结构粒度上操作,且将键信息视为辅助边属性而非独立语义层。本文提出MolCHG,一种基于新型组合分层图(Compositional Hierarchical Graph)的多层级自监督预训练框架,将分子结构组织为三个语义层级上的四类节点。通过引入与原子图并行的键图,使键级信息获得独立演化的节点表示,让片段节点能平等聚合原子级与键级语义。设计三种层级特定的预训练目标:片段内原子-键跨视图对比任务以对齐原子视图与键视图表示;片段级官能团预测任务注入领域化学知识;图级结构预测任务编码全局分子拓扑。在九个MoleculeNet基准测试中,MolCHG在七个数据集上取得最优表现,分类与回归任务均显著领先,其余任务也保持与最强基线相当的竞争力。消融实验进一步验证多层级监督信号具有互补性,各组件均对整体性能有贡献。
原文摘要 · Abstract (English)
Self-supervised pretraining on molecular graphs has emerged as a promising approach for molecular property prediction, yet most existing methods operate at a single structural granularity and treat bond information as auxiliary edge attributes rather than as an independent semantic layer. In this work, we propose MolCHG, a multi-level self-supervised pretraining framework built upon a novel Compositional Hierarchical Graph that organizes molecular structure into four types of nodes across three semantic levels. By introducing a bond graph that operates in parallel with the atom graph, our architecture elevates bond-level information to independently evolving node representations, enabling fragment nodes to aggregate atom-level and bond-level semantics on an equal footing. We design three level-specific pretraining objectives: an atom-bond cross-view contrastive task that aligns the atom-view and bond-view representations within each fragment, a fragment-level functional group prediction task to inject domain-relevant chemical knowledge, and graph-level structure prediction tasks to encode global molecular topology. Experiments on nine MoleculeNet benchmarks demonstrate that MolCHG achieves the best performance on seven datasets across both classification and regression tasks, remaining competitive with the strongest baselines on the rest. Ablation studies further confirm that the multi-level supervision signals are complementary and that each component contributes to the overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。