测试机器学习势能模型对未见分子的组合泛化能力
Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials

- 设计四个需组合泛化的基准任务,检验模型能否掌握化学结构规律
- 最先进模型在分布外数据上误差比分布内高一个数量级
- 适合关注模型可解释性与物理规律学习的研究者
机器学习原子间势能在计算化学和材料科学中具有基础作用,支持从分子动力学模拟到药物设计和材料发现的应用。尽管近期方法能高精度预测原子间力,但其对未见分子的泛化能力仍不明确:是学习了化学的组合结构(即分子片段及其组合如何决定性质),还是仅依赖训练样本中的模式进行插值?为回答此问题,我们提出一个包含四个任务的基准测试,要求模型在训练中未见过的分子上实现组合泛化。每个任务中,训练数据被精心设计,使得若模型理解物理原理,则泛化应可行。实验分析表明,这些任务对当前最先进的模型极具挑战性,即使使用在数百万分子上预训练的基础模型,分布外测试的误差也常比分布内高出一个数量级。
原文摘要 · Abstract (English)
Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular dynamics simulations to drug design and materials discovery. While recent approaches can estimate inter-atomic forces with high precision, it remains unclear to what extent they can generalise to previously unseen molecules. Do they learn the compositional structure of chemistry, capturing how molecular fragments and their combinations determine properties, or do they primarily learn to interpolate patterns that are specific to the training examples? To address this question, we propose a benchmark consisting of four tasks that require some form of compositional generalisation. In each task, models are tested on molecules that were unseen during training, but the training data is chosen such that generalisation to the test examples should be feasible for models that learn the underlying physical principles. Our empirical analysis shows that the considered tasks are highly challenging for state-of-the-art models, with errors on out-of-distribution examples often an order of magnitude higher than on in-distribution examples, even when using foundation models that have been pre-trained on millions of molecules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。