arXiv:2605.13788cs.LG2026-05被引 1

提出力感知核方法,实现高效精准的机器学习势能主动学习。

Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs

论文配图:Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
图 1 · 摘自论文原文
  • 用分块后验方差筛选实现线性扩展,支持数万结构快速筛选。
  • 引入力感知神经正切核,在OC20上达到最低能量与力误差。
  • 对分布偏移鲁棒,适合大池候选集和基础模型微调场景。

机器学习原子势能(MLIPs)的主动学习需应对大规模候选集、利用能量-力联合监督及分布偏移下的鲁棒性挑战。本文提出一种基于分块特征空间后验方差筛选的线性扩展获取框架,避免显式计算候选集与训练集核矩阵,可在数小时内完成约20万结构的筛选,适用于基于分子相似性的各类策略。通过混合参数-坐标导数,将神经正切核(NTK)扩展至力感知设置,得到力NTK与联合能量-力NTK,提供向量场预测的自然相似性度量。在OC20数据集上,联合能量-力NTK表现最优,各项指标均低于其他方法。在T1x、PMechDB和RGD基准测试中,力NTK方法性能媲美现有基线,且远优于委员会法的效率。在T1x上的受控偏移实验显示,基于预训练模型嵌入与NTK的获取方法保持稳定,而委员会法方差更高。结果表明,单一预训练MLIP可实现可扩展、力感知且分布鲁棒的主动学习,适用于基础模型微调。

原文摘要 · Abstract (English)

Active learning for machine-learning interatomic potentials (MLIPs) must address several challenges to be practical: scaling to large candidate pools, leveraging energy-force supervision, and maintaining robustness when candidate pools are biased relative to the target distribution. In this work, we jointly address these challenges. We first introduce a linearly scaling acquisition framework based on chunked feature-space posterior-variance shortlisting. By avoiding materialisation of the candidate and train set kernels, this approach enables screening of ~200k structures within hours and applies broadly to acquisition strategies that score candidates based on molecular similarity metrics. We then extend the Neural Tangent Kernel (NTK) to a force-aware setting via mixed parameter-coordinate derivatives, yielding a force NTK and a joint energy-force NTK that provide natural similarity metrics for vector-field prediction. We demonstrate the effectiveness of the joint energy-force NTK on the OC20 dataset, where force-aware acquisition is crucial: it achieves the lowest energy and force MAE and RMSE across all metrics and distribution splits. Across T1x, PMechDB, and RGD benchmarks, our force NTK methods remain competitive with established baselines while being significantly more efficient than committee-based approaches. Under a controlled candidate-pool shift case study on T1x, acquisition based on pretrained MLIP embeddings and NTKs remains robust, whereas committee-based methods exhibit higher variance. Overall, these results show that a single pretrained MLIP can enable scalable, force-aware, and distribution-robust active learning for foundation-model fine-tuning.

主动学习机器学习势能神经正切核力感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。