arXiv:2508.15577physics.chem-phcs.AI2025-08被引 3

用低精度数据降低偏差,让化学构型主动学习更高效

LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space

  • 用低成本低精度数据近似模型偏差,替代传统方差最小化
  • 在量子化学任务中减少训练数据消耗达十倍以上
  • 适合需要高效数据利用的分子性质预测场景

主动学习有望在构建机器学习模型时实现最优样本选择。传统方法通常通过最小化模型方差来降低预测误差,但实际效率常不如随机采样。基于偏差-方差分解,我们提出应最小化模型偏差而非方差。该方法利用低成本的低精度数据(如Δ-ML或多精度学习中的信息),在量子化学多个应用中验证,包括激发能与从头算势能面预测。实验表明,相比标准主动学习,该方法可将训练数据需求减少近一个数量级。

原文摘要 · Abstract (English)

Active learning promises to provide an optimal training sample selection procedure in the construction of machine learning models. It often relies on minimizing the model's variance, which is assumed to decrease the prediction error. Still, it is frequently even less efficient than pure random sampling. Motivated by the bias-variance decomposition, we propose to minimize the model's bias instead of its variance. By doing so, we are able to almost exactly match the best-case error over all possible greedy sample selection procedures for a relevant application. Our bias approximation is based on using cheap to calculate low fidelity data as known from $Δ$-ML or multifidelity machine learning. We exemplify our approach for a wider class of applications in quantum chemistry including predicting excitation energies and ab initio potential energy surfaces. Here, the proposed method reduces training data consumption by up to an order of magnitude compared to standard active learning.

主动学习量子化学多精度学习数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。