arXiv:2603.17196cs.LG2026-03被引 3

用自条件去噪提升原子级表示学习,跨多种材料类型表现卓越。

Self-Conditioned Denoising for Atomistic Representation Learning

  • 通过自嵌入条件去噪,实现跨原子数据域的统一预训练
  • 小模型经SCD预训练后性能超越更大规模的基线模型
  • 适用于分子、蛋白、周期性材料等多领域,尤其适合非平衡构型

大模型在自然语言处理和计算机视觉中的成功,推动了物理科学领域基础模型的发展。然而,基于原子数据的预训练策略仍不充分。目前,基于密度泛函理论力-能标签的大规模监督预训练效果最佳,显著优于现有自监督学习方法,后者仅限于基态构型或单一数据域。本文提出自条件去噪(SCD),一种与主干网络无关的重建目标,利用自嵌入对任意原子数据域(包括小分子、蛋白质、周期性材料及非平衡构型)进行条件去噪。在控制主干结构与预训练数据集的前提下,SCD在下游任务中显著优于以往自监督方法,并达到或超过监督力-能预训练的性能。结果显示,一个小型快速的图神经网络经由SCD预训练后,在多个任务和领域中表现可媲美甚至超越更大模型在更大规模有标签或无标签数据上的预训练结果。代码已开源:https://github.com/TyJPerez/SelfConditionedDenoisingAtoms

原文摘要 · Abstract (English)

The success of large-scale pretraining in NLP and computer vision has catalyzed growing efforts to develop analogous foundation models for the physical sciences. However, pretraining strategies using atomistic data remain underexplored. To date, large-scale supervised pretraining on DFT force-energy labels has provided the strongest performance gains to downstream property prediction, out-performing existing methods of self-supervised learning (SSL) which remain limited to ground-state geometries, and/or single domains of atomistic data. We address these shortcomings with Self-Conditioned Denoising (SCD), a backbone-agnostic reconstruction objective that utilizes self-embeddings for conditional denoising across any domain of atomistic data, including small molecules, proteins, periodic materials, and 'non-equilibrium' geometries. When controlled for backbone architecture and pretraining dataset, SCD significantly outperforms previous SSL methods on downstream benchmarks and matches or exceeds the performance of supervised force-energy pretraining. We show that a small, fast GNN pretrained by SCD can achieve competitive or superior performance to larger models pretrained on significantly larger labeled or unlabeled datasets, across tasks in multiple domains. Our code is available at: https://github.com/TyJPerez/SelfConditionedDenoisingAtoms

原子表示学习自监督学习图神经网络材料科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。