用分类方法评估机器学习势函数的可靠性,更准且可扩展。
Knowing when to trust machine-learned interatomic potentials

- 用轻量判别器分析预训练模型的原子特征,转为可靠性分类
- 可靠性概率与真实误差单调相关,优于传统集成分歧方法
- 无需修改模型,可解释性好,适合大模型和多架构部署
现有机器学习势函数(MLIP)不确定性量化方法依赖独立训练的模型集成,但其规模随基础型模型增长而恶化,且成员间分歧信号与单分子预测误差相关性弱。本文提出一种后处理方法PROBE(Post-hoc Reliability frOm Backbone Embeddings),通过一个紧凑的判别分类器分析预训练MLIP的冻结原子表示,将不确定性量化重构为选择性分类问题。该方法输出每个预测的可靠性概率,与实际误差单调一致,且不需修改原模型。在多个大规模保留测试集及两种结构不同的MLIP架构上,PROBE均优于基于集成分歧的二元可靠性信号,且随着骨干表示能力增强,性能持续提升,展现出向基础规模模型演进的有利可扩展性。多头自注意力还生成原子级重要性图,提供化学可解释诊断,且无额外计算开销。PROBE为后处理、架构无关,可直接部署于任意暴露原子表示的MLIP。
原文摘要 · Abstract (English)
Prevailing machine-learned interatomic potential (MLIP) uncertainty-quantification methods rely on ensembles of independently trained backbones. These methods scale unfavorably with foundation-scale MLIPs, and their member-disagreement signals correlate weakly with per-molecule prediction error. Here we probe the frozen per-atom representations of a pretrained MLIP with a compact discriminative classifier, recasting MLIP uncertainty quantification as selective classification rather than error regression. The resulting method, PROBE (Post-hoc Reliability frOm Backbone Embeddings), produces a per-prediction reliability probability that monotonically tracks actual error without modification to the underlying model. Across large held-out evaluation sets and two structurally distinct MLIP architectures, PROBE outperforms ensemble disagreement as a binary reliability signal, which strengthens with the expressiveness of the backbone representation, implying a favorable scaling trajectory toward foundation-scale MLIPs. Multi-head self-attention additionally yields per-atom importance maps, providing chemically interpretable diagnostics at no additional computational cost. PROBE is post-hoc and architecture-agnostic, and is directly deployable on any MLIP that exposes per-atom representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。