arXiv:2503.08674cs.LGcond-mat.mtrl-sci2025-03被引 10

提出两种轻量测试优化方法,提升分子力场模型在分布外数据的表现。

Understanding and Mitigating Distribution Shifts For Machine Learning Force Fields

  • 基于图谱理论调整测试图结构,使其更贴近训练分布。
  • 利用廉价物理先验在测试时进行梯度优化,改善分布外系统表征。
  • 无需昂贵参考数据,计算开销小,适合实际部署场景。

机器学习力场(MLFF)是替代昂贵的从头量子力学分子模拟的有前途方案。由于关注的化学空间多样且新数据生成成本高,理解MLFF在训练分布外的泛化能力至关重要。本文通过诊断实验分析化学数据集,揭示了即使在大规模基础模型上也存在显著的分布偏移挑战。我们推测当前监督训练方法对MLFF正则化不足,导致过拟合和对分布外系统表征不佳。为此,提出两种新的测试时优化策略:第一种基于谱图理论,调整测试图的边以匹配训练时的图结构;第二种通过使用廉价物理先验的辅助目标,在测试时进行梯度更新,改进分布外系统的表示。这两种策略均在不依赖昂贵从头计算标签的前提下,显著降低了分布外系统的误差,表明MLFF具备建模多样化化学空间的能力,但当前训练方式未能有效释放这种潜力。实验建立了下一代MLFF泛化能力评估的明确基准。代码已公开于 https://tkreiman.github.io/projects/mlff_distribution_shifts/。

原文摘要 · Abstract (English)

Machine Learning Force Fields (MLFFs) are a promising alternative to expensive ab initio quantum mechanical molecular simulations. Given the diversity of chemical spaces that are of interest and the cost of generating new data, it is important to understand how MLFFs generalize beyond their training distributions. In order to characterize and better understand distribution shifts in MLFFs, we conduct diagnostic experiments on chemical datasets, revealing common shifts that pose significant challenges, even for large foundation models trained on extensive data. Based on these observations, we hypothesize that current supervised training methods inadequately regularize MLFFs, resulting in overfitting and learning poor representations of out-of-distribution systems. We then propose two new methods as initial steps for mitigating distribution shifts for MLFFs. Our methods focus on test-time refinement strategies that incur minimal computational cost and do not use expensive ab initio reference labels. The first strategy, based on spectral graph theory, modifies the edges of test graphs to align with graph structures seen during training. Our second strategy improves representations for out-of-distribution systems at test-time by taking gradient steps using an auxiliary objective, such as a cheap physical prior. Our test-time refinement strategies significantly reduce errors on out-of-distribution systems, suggesting that MLFFs are capable of and can move towards modeling diverse chemical spaces, but are not being effectively trained to do so. Our experiments establish clear benchmarks for evaluating the generalization capabilities of the next generation of MLFFs. Our code is available at https://tkreiman.github.io/projects/mlff_distribution_shifts/.

分子模拟分布外泛化力场模型测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。