用无监督图自编码器自动发现原子结构数据中的相似区域,消除采样偏差。
Unsupervised Atomic Data Mining via Multi-Kernel Graph Autoencoders for Machine Learning Force Fields
- 基于多核注意力机制的图自编码器,捕捉原子构型几何敏感性。
- 在铌、钽、铁数据集上实现高效聚类,显著减少采样偏差。
- 无需标签即可优化数据集,适合材料力场训练前的数据清理。
构建化学多样性高且无采样偏差的数据集对训练高效通用的分子力场至关重要。然而,计算化学与材料科学中常见的数据生成方法易集中在势能面的某些区域,这些区域难以识别且常与人类直觉不符,导致系统性偏差难以消除。传统聚类与降采样方法虽有一定作用,但受限于原子描述符的高维特性,易造成信息丢失或无法准确区分不同区域。本文提出多核边注意力图自编码器(MEAGraph),一种无监督原子数据挖掘方法。该模型结合多重线性核变换与注意力消息传递,有效捕捉原子环境的几何敏感性,实现无需标签的高效数据聚类与精简。在铌、钽、铁数据集上的实验表明,该方法能精准分组相似原子环境,支持基础剪枝策略以去除采样偏差。该方法为表示学习、聚类分析、异常检测与数据集优化提供了有效工具。
原文摘要 · Abstract (English)
Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。