为机器学习势能提供可验证的安全证书,提升材料筛选可靠性
Proof-Carrying Materials: Falsifiable Safety Certificates for Machine-Learned Interatomic Potentials
- 用对抗性验证+置信区间优化+形式化证明三阶段构建安全证书
- 单模型筛选漏检93%稳定材料,新方法发现62个新增稳定物质量
- 适用于需要高可靠性的材料筛选场景,尤其适合多模型交叉验证
机器学习的原子间势能(MLIPs)在高通量材料筛选中广泛应用,但缺乏正式的可靠性保障。我们发现,仅使用一个MLIP作为稳定性过滤器时,在2.5万种材料的基准测试中,会漏掉93%的密度泛函理论(DFT)稳定的材料(召回率0.07)。为此,提出「可携带证明材料」(Proof-Carrying Materials, PCM),通过三个阶段:跨组分空间的对抗性证伪、95%置信区间的自举包络优化,以及用Lean 4进行形式化认证。对CHGNet、TensorNet和MACE的审计揭示了各架构特有的盲区,成对误差相关性极低(r ≤ 0.13;n = 5,000),并通过独立的Quantum ESPRESSO验证确认(20/20收敛;中位数DFT/CHGNet力比值为12倍)。基于PCM发现特征训练的风险模型,在未见材料上预测失败表现优异(AUC-ROC = 0.938 ± 0.004),且跨模型迁移有效(跨模型AUC-ROC ≈ 0.70;特征重要性相关系数r = 0.877)。在热电材料筛选案例中,PCM审计协议比单模型筛选多发现62种稳定材料,发现率提升25%。
原文摘要 · Abstract (English)
Machine-learned interatomic potentials (MLIPs) are deployed for high-throughput materials screening without formal reliability guarantees. We show that a single MLIP used as a stability filter misses 93% of density functional theory (DFT)-stable materials (recall 0.07) on a 25,000-material benchmark. Proof-Carrying Materials (PCM) closes this gap through three stages: adversarial falsification across compositional space, bootstrap envelope refinement with 95% confidence intervals, and Lean 4 formal certification. Auditing CHGNet, TensorNet and MACE reveals architecture-specific blind spots with near-zero pairwise error correlations (r <= 0.13; n = 5,000), confirmed by independent Quantum ESPRESSO validation (20/20 converged; median DFT/CHGNet force ratio 12x). A risk model trained on PCM-discovered features predicts failures on unseen materials (AUC-ROC = 0.938 +/- 0.004) and transfers across architectures (cross-MLIP AUC-ROC ~ 0.70; feature importance r = 0.877). In a thermoelectric screening case study, PCM-audited protocols discover 62 additional stable materials missed by single-MLIP screening - a 25% improvement in discovery yield.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。