arXiv:2410.16122physics.chem-phcs.LG2024-10

用整数线性规划优化分子数据集选择,提升大分子性质预测精度。

Integer linear programming for unsupervised training set selection in molecular machine learning

  • 基于原子级局部相似性构建整数线性规划模型,自动筛选最优训练集。
  • 在预测大于训练集范围的大分子时,性能显著优于现有无监督方法。
  • 适合需要高泛化能力的化学机器学习研究者使用。

整数线性规划(ILP)是一种通过整数决策变量自然描述的线性优化方法。在面向化学的物理启发式机器学习中,我们展示了将ILP用于选择分子训练集以预测广义尺度性质的有效性。实验表明,该算法在预测超出训练集分子规模的化合物性质时,显著优于现有的无监督训练集选择方法。性能提升源于基于局部相似性(即原子级)的选样策略,以及一种能高效求解最优解的独特ILP框架。本工作提供了一种实用算法,可提升物理启发式机器学习模型的表现,并揭示了其与现有方法的本质差异。

原文摘要 · Abstract (English)

Integer linear programming (ILP) is an elegant approach to solve linear optimization problems, naturally described using integer decision variables. Within the context of physics-inspired machine learning applied to chemistry, we demonstrate the relevance of an ILP formulation to select molecular training sets for predictions of size-extensive properties. We show that our algorithm outperforms existing unsupervised training set selection approaches, especially when predicting properties of molecules larger than those present in the training set. We argue that the reason for the improved performance is due to the selection that is based on the notion of local similarity (i.e., per-atom) and a unique ILP approach that finds optimal solutions efficiently. Altogether, this work provides a practical algorithm to improve the performance of physics-inspired machine learning models and offers insights into the conceptual differences with existing training set selection approaches.

分子机器学习整数规划训练集选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。