arXiv:2512.07064cs.LGcs.AI2025-12被引 1

揭秘分子图自监督学习中掩码设计的真正关键

Self-Supervised Learning on Molecular Graphs: A Systematic Investigation of Masking Design

  • 构建统一概率框架,系统对比掩码策略
  • 均匀采样比复杂分布更有效,目标设计更重要
  • 语义丰富的目标+图变压器架构提升显著

自监督学习(SSL)在分子表示学习中占据核心地位。然而,许多基于掩码的预训练创新仅作为启发式方法引入,缺乏严谨评估,导致难以判断哪些设计选择真正有效。本文将整个预训练-微调流程纳入统一的概率框架,实现对掩码策略的透明比较与深入理解。在此基础上,我们在严格控制条件下系统研究了三个核心设计维度:掩码分布、预测目标与编码器架构。进一步利用信息论度量评估预训练信号的有用性,并将其与下游任务性能关联。研究发现:对于常见的节点级预测任务,复杂掩码分布并未优于均匀采样;相反,预测目标的选择及其与编码器架构的协同效应更为关键。具体而言,采用语义更丰富的预测目标可带来显著的下游性能提升,尤其在搭配表达能力强的图变换器编码器时效果更佳。这些发现为开发更有效的分子图自监督学习方法提供了实用指导。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) plays a central role in molecular representation learning. Yet, many recent innovations in masking-based pretraining are introduced as heuristics and lack principled evaluation, obscuring which design choices are genuinely effective. This work cast the entire pretrain-finetune workflow into a unified probabilistic framework, enabling a transparent comparison and deeper understanding of masking strategies. Building on this formalism, we conduct a controlled study of three core design dimensions: masking distribution, prediction target, and encoder architecture, under rigorously controlled settings. We further employ information-theoretic measures to assess the informativeness of pretraining signals and connect them to empirically benchmarked downstream performance. Our findings reveal a surprising insight: sophisticated masking distributions offer no consistent benefit over uniform sampling for common node-level prediction tasks. Instead, the choice of prediction target and its synergy with the encoder architecture are far more critical. Specifically, shifting to semantically richer targets yields substantial downstream improvements, particularly when paired with expressive Graph Transformer encoders. These insights offer practical guidance for developing more effective SSL methods for molecular graphs.

自监督学习分子图掩码设计图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。