arXiv:2506.21788cs.LGcond-mat.mtrl-sci2025-06被引 1

用多任务并行加速图模型预训练,处理海量原子数据更高效。

Multi-task parallelism for robust pre-training of graph foundation models on multi-source, multi-fidelity atomistic modeling data

  • 将多任务解码头分散到多个GPU并行计算,提升训练效率。
  • 在超2400万结构上训练,跨三台异构超算实现良好扩展性。
  • 适合需要大规模原子建模的科研人员和高性能计算团队。

基于图神经网络的图基础模型有望实现可持续、高效的原子级建模。为应对预训练阶段处理多源、多保真度数据的挑战,现有研究采用多任务学习:共享的消息传递层先处理输入的原子结构(不区分来源),再路由至多个解码头以预测特定数据输出。该方法稳定了预训练过程,并提升了模型对未探索化学区域的迁移能力。初步在约四百万个结构上的结果令人鼓舞,但其在更大、更多样化数据集上的泛化能力及在超级计算机上的可扩展性仍存疑问。本文提出一种多任务并行方法,通过GPU加速将每个解码头分布到计算资源上。该方法在开源HydraGNN架构中实现,训练数据来自五个数据集,超过2400万结构,测试于Perlmutter、Aurora和Frontier超算,均展现出良好的扩展性,验证了其在三种高度异构超级计算架构上的高效性。

原文摘要 · Abstract (English)

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model's transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

图神经网络原子建模多任务学习超算加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。