用物理启发搜索发现机器学习势能模型的隐藏故障,超越传统基准测试。
MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials

- 基于物理规律主动生成可验证的故障假设
- 发现MACE-MPA-0在含氧/氟吸附体系中能量预测异常
- 适合关注模型可靠性的材料模拟研究者
通用机器学习原子间势(u-MLIPs)旨在跨多种构型泛化。基准测试虽可重复评估,但可能无法暴露其预设范围外的失效。本文提出MLIP Detective,一种智能体框架,通过物理启发搜索补足基准评估,主动发现隐藏故障模式。该框架从基准证据出发,生成可证伪的物理驱动故障假设,以低成本模拟筛选,仅将最可疑案例提交人类专家并附验证方案。无需特定任务提示,该系统成功识别并表征了MACE-MPA-0中的系统性异常:在部分松弛的吸附剂-表面体系中,含氧或氟的吸附物能量被高估,甚至高于分离碎片总能。通过跨模型对比,进一步推断该异常可能源于训练数据,与近期报告一致。
原文摘要 · Abstract (English)
Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark evidence, MLIP Detective generates falsifiable, physics-informed failure hypotheses, screens them with inexpensive simulations, and escalates only the most suspicious cases to human experts together with proposed verification protocols. Without issue-specific prompting, MLIP Detective identified and characterized a systematic anomaly in MACE-MPA-0: the model predicted some relaxed adsorbate-surface systems involving O- or F-containing adsorbates to be higher in energy than their corresponding separated fragments. Using cross-model comparisons, MLIP Detective further inferred a likely training-data origin for the anomaly, consistent with recent reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。