arXiv:2609.05877cs.LG2026-09

根据数据量大小,选择结构多样性或模型分歧来压缩原子势训练集。

A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials

  • 按预算动态选择:小数据用结构覆盖,大数据用模型分歧
  • 20%数据时,分歧法比覆盖法降低0.46%~5.89%的力预测误差
  • 适合需要高效压缩已有标注数据的材料模拟研究者

构建机器学习原子势的紧凑训练集需权衡保留结构多样性或聚焦模型预测分歧。本文通过预算解析比较在GAP-20 Carbon和合并修订版MD17数据集上重训练的MACE模型表现。以结构覆盖与响应引导选择器(基于覆盖模型与全数据参考的分歧)对比,发现5%数据下覆盖法绝对误差低于随机采样;而在1%和5%时响应法误差更大,但20%时反转——此时响应法使独立持有测试力误差降低0.46%~5.89%,八组训练种子均优于覆盖法,六项误差低于全数据基准。平均力误差减少0.164~0.167 meV Å⁻¹,尾部及掩码端点收益更显著。互补分析表明,学习相似性保持覆盖排序,而固定模型误差选择法误差更高。结果确立数据保留预算为关键决策变量,并直接验证响应引导压缩在何种条件下优于结构覆盖。

原文摘要 · Abstract (English)

Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. The better choice can depend on how much data is retained, making a comparison at one training-set size insufficient. Here we link selection criteria to prediction accuracy through a budget-resolved comparison of retrained MACE models on GAP-20 Carbon and pooled revised MD17. Structural coverage is compared with a response-guided selector that targets disagreement between a coverage-trained model and a full-data reference. This retrospective response witness tests the value of model disagreement for compressing an already labelled pool. At 5\%, coverage gives smaller absolute deviations from the full-data error than random sampling across four force endpoints in both datasets. The witness has larger deviations than coverage at 1\% and 5\%, but the ordering reverses at 20\%. At 20\%, witness-selected models also lower direct held-out force errors by 0.46--5.89\% relative to coverage, with all eight paired training-seed intervals favouring the witness. Six errors fall below the full-data reference. Mean force-error reductions are 0.164--0.167~meV~$\text{\AA}^{-1}$, with larger gains for tail and masked endpoints. Complementary analyses show that learned similarity preserves the coverage ranking, while selecting by frozen-model error gives higher error than embedding coverage. These findings establish retained-data budget as a deciding variable in atomistic training-set selection and provide a direct test of when response-guided compression improves on structural coverage.

原子势训练集选择机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。