arXiv:2503.23515physics.chem-phcs.CV2025-03被引 4

构建原子环境描述符的最小完备集,提升机器学习模型效率与表达力。

Optimal Invariant Bases for Atomistic Machine Learning

  • 基于模式识别技术剔除冗余描述符,得到最小完备集。
  • 新架构可识别至五体相互作用,基准测试表现优异且计算开销低。
  • 适用于需要高效高精度原子建模的研究者,如材料设计与分子模拟。

原子构型的机器学习表示已催生多种描述符,用于刻画原子局部环境。然而,许多现有描述符存在不完整或功能依赖问题:不完整集无法表达所有有意义的环境变化;而完全构造常因高度功能依赖导致部分描述符可由其他描述符函数表示,造成冗余,降低区分能力并增加计算负担。本文借鉴模式识别方法,从已有原子描述符中移除函数依赖项,获得满足完备性的最小集合。首先,改进现有的原子簇展开(Atomistic Cluster Expansion),得到更高效的描述符子集;其次,将该最小完备集应用于基于标量神经网络的不完整构造,提出一种新型消息传递网络架构,每个神经元可识别最多五体相互作用,借助最优笛卡尔张量不变量实现高表达性。该架构在多个前沿基准测试中表现出色,同时保持低计算成本。研究成果不仅提升了模型性能,也为各类应用提供了兼具低复杂度与高表达力的不变基范式。

原文摘要 · Abstract (English)

The representation of atomic configurations for machine learning models has led to the development of numerous descriptors, often to describe the local environment of atoms. However, many of these representations are incomplete and/or functionally dependent. Incomplete descriptor sets are unable to represent all meaningful changes in the atomic environment. Complete constructions of atomic environment descriptors, on the other hand, often suffer from a high degree of functional dependence, where some descriptors can be written as functions of the others. These redundant descriptors do not provide additional power to discriminate between different atomic environments and increase the computational burden. By employing techniques from the pattern recognition literature to existing atomistic representations, we remove descriptors that are functions of other descriptors to produce the smallest possible set that satisfies completeness. We apply this in two ways: first we refine an existing description, the Atomistic Cluster Expansion. We show that this yields a more efficient subset of descriptors. Second, we augment an incomplete construction based on a scalar neural network, yielding a new message-passing network architecture that can recognize up to 5-body patterns in each neuron by taking advantage of an optimal set of Cartesian tensor invariants. This architecture shows strong accuracy on state-of-the-art benchmarks while retaining low computational cost. Our results not only yield improved models, but point the way to classes of invariant bases that minimize cost while maximizing expressivity for a host of applications.

原子建模不变基神经网络材料科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。