arXiv:2501.09009physics.chem-phcond-mat.mtrl-sci2025-01ICLR被引 40

用能量赫斯矩阵蒸馏大模型,打造更快更准的专用分子力场。

Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians

  • 以教师模型的能量赫斯矩阵指导学生模型训练,实现知识高效迁移。
  • 专用力场推理速度最快提升20倍,精度不降反升。
  • 适合需要快速模拟特定化学体系的研究者,兼顾速度与物理一致性。

基础模型(FM)正推动机器学习力场(MLFF)发展,利用通用表征和可扩展训练完成多种计算化学任务。尽管MLFF基础模型已接近第一性原理方法的精度,但推理速度仍需提升。此外,研究虽聚焦跨化学空间的通用模型,实际应用中用户通常仅关注特定小范围体系。这凸显了对快速、专用化MLFF的需求:既保持测试时的物理合理性,又具备训练时的可扩展性。本文提出一种将通用表征从MLFF基础模型蒸馏至更小、更快的专用模型的方法。该方法基于知识蒸馏框架,使小型“学生”模型通过匹配“教师”模型的能量赫斯矩阵来学习。所得专用力场比原基础模型快达20倍,同时保留甚至超越其性能及未蒸馏模型的表现。我们还证明,将具有直接力参数化的教师模型蒸馏至以保守力(即势能导数计算)训练的学生模型,能有效利用大规模教师模型的表征,提升精度,并在测试时分子动力学模拟中保持能量守恒。本工作提出了新的MLFF开发范式:发布基础模型的同时,配套提供针对常见化学子集的轻量级模拟‘引擎’。

原文摘要 · Abstract (English)

The foundation model (FM) paradigm is transforming Machine Learning Force Fields (MLFFs), leveraging general-purpose representations and scalable training to perform a variety of computational chemistry tasks. Although MLFF FMs have begun to close the accuracy gap relative to first-principles methods, there is still a strong need for faster inference speed. Additionally, while research is increasingly focused on general-purpose models which transfer across chemical space, practitioners typically only study a small subset of systems at a given time. This underscores the need for fast, specialized MLFFs relevant to specific downstream applications, which preserve test-time physical soundness while maintaining train-time scalability. In this work, we introduce a method for transferring general-purpose representations from MLFF foundation models to smaller, faster MLFFs specialized to specific regions of chemical space. We formulate our approach as a knowledge distillation procedure, where the smaller "student" MLFF is trained to match the Hessians of the energy predictions of the "teacher" foundation model. Our specialized MLFFs can be up to 20 $\times$ faster than the original foundation model, while retaining, and in some cases exceeding, its performance and that of undistilled models. We also show that distilling from a teacher model with a direct force parameterization into a student model trained with conservative forces (i.e., computed as derivatives of the potential energy) successfully leverages the representations from the large-scale teacher for improved accuracy, while maintaining energy conservation during test-time molecular dynamics simulations. More broadly, our work suggests a new paradigm for MLFF development, in which foundation models are released along with smaller, specialized simulation "engines" for common chemical subsets.

机器学习力场知识蒸馏分子模拟加速计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。