用预训练提升全原子图神经网络泛化能力,实现零样本迁移。
Pushing the Limits of All-Atom Geometric Graph Neural Networks: Pre-Training, Scaling and Zero-Shot Transfer
- 采用自监督预训练增强几何图神经网络表达能力
- 发现预训练阶段无幂律缩放,但能有效提升零样本迁移性能
- 适合药物设计与分子性质预测等需要强泛化的场景
构建可迁移的分子与生物系统构象表征描述符在药物发现、基于学习的分子动力学和蛋白质机制分析中具有广泛应用。包含全原子信息的几何图神经网络(Geom-GNN)通过作为下游任务的通用可学习几何描述符,已革新原子级模拟,用于预测原子间势能和分子性质。然而,现有方法通常在特定下游任务上进行监督训练,受限于高质量数据缺乏和标签不准确,导致泛化能力差,在分布外(OOD)场景下性能下降。本文探索使用预训练的Geom-GNN作为可迁移且高效的几何描述符以提升泛化能力。研究了不同学习设置下Geom-GNN的缩放行为,包括自监督预训练、监督与无监督学习。发现不同架构在预训练任务上的表达能力存在差异;有趣的是,Geom-GNN在预训练任务中不遵循幂律缩放,且在具有量子化学标签的监督任务中普遍缺乏可预测的缩放规律,这对新分子筛选与设计至关重要。更重要的是,我们展示了如何将全原子图嵌入与其它神经网络架构有机融合以增强表达能力;同时,潜在空间的低维投影与传统几何描述符高度一致。
原文摘要 · Abstract (English)
Constructing transferable descriptors for conformation representation of molecular and biological systems finds numerous applications in drug discovery, learning-based molecular dynamics, and protein mechanism analysis. Geometric graph neural networks (Geom-GNNs) with all-atom information have transformed atomistic simulations by serving as a general learnable geometric descriptors for downstream tasks including prediction of interatomic potential and molecular properties. However, common practices involve supervising Geom-GNNs on specific downstream tasks, which suffer from the lack of high-quality data and inaccurate labels leading to poor generalization and performance degradation on out-of-distribution (OOD) scenarios. In this work, we explored the possibility of using pre-trained Geom-GNNs as transferable and highly effective geometric descriptors for improved generalization. To explore their representation power, we studied the scaling behaviors of Geom-GNNs under self-supervised pre-training, supervised and unsupervised learning setups. We find that the expressive power of different architectures can differ on the pre-training task. Interestingly, Geom-GNNs do not follow the power-law scaling on the pre-training task, and universally lack predictable scaling behavior on the supervised tasks with quantum chemical labels important for screening and design of novel molecules. More importantly, we demonstrate how all-atom graph embedding can be organically combined with other neural architectures to enhance the expressive power. Meanwhile, the low-dimensional projection of the latent space shows excellent agreement with conventional geometrical descriptors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。