用少标签实现全数据精度,让机器学习力场更高效可靠
Full-data accuracy with fewer labels for training and fine-tuning machine-learning force fields

- 基于最后一层投影回归的快速不确定性估计方法
- 仅用少量量子力学标签即可达到全数据训练精度
- 适合需要高效标注的分子、电解质等系统建模场景
机器学习力场(MLFF)的可靠性受限于训练分布,构建多样化训练集是训练和基础模型微调的核心瓶颈。传统主动学习依赖模型委员会不确定性,但对基础模型而言每次评估需独立微调,成本过高。本文提出基于最后层投影回归(LLPR)的主动学习流程,该方法可在单次前向传播中快速估算每构型的不确定性。在分子、凝聚相及电解质系统中,LLPR识别出紧凑且高价值的训练集,仅使用少量电子结构标签即可恢复全数据精度。在基础模型微调中,LLPR选择的样本以远少于随机采样的标签数达到全池微调上限。在迭代电解质微调中,LLPR可提前检测非物理解构,提供绝对力误差阈值,并实现学习循环的自动终止。最终模型准确再现参考密度与离子配位结构,为各类MLFF训练提供可扩展的不确定性量化策略。
原文摘要 · Abstract (English)
Machine-learning force fields (MLFFs) are reliable only near their training distribution, making efficient construction of diverse training sets a major bottleneck for both train-from-scratch and foundation fine-tuning workflows. Active learning can reduce this cost, but standard model-committee uncertainty is impractical for foundation MLFFs because each committee member requires a separate fine-tuning run. We present an active-learning workflow based on last-layer-projection regression (LLPR), a forward-pass-cheap per-configuration uncertainty estimator. Across molecular, condensed-phase, and electrolyte systems, LLPR identifies compact, high-value training sets that recover full-data accuracy using only a small fraction of electronic-structure labels. In foundation-model fine-tuning, LLPR-selected configurations reach the full-pool fine-tuning ceiling with substantially fewer labels than random selection. In iterative electrolyte fine-tuning, LLPR detects unphysical local coordination before DFT labelling, provides an absolute force-error threshold, and enables automatic termination of the learning loop. The resulting models reproduce reference density and ion-coordination structure, providing a scalable uncertainty-quantification strategy across MLFF training regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。