arXiv:2506.11100cs.LGcs.AI2025-06

用主动学习减少中子衍射结构预测的训练数据量,提速并提准。

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

  • 基于不确定性采样的主动学习策略,优先生成模型最不确定的数据。
  • 仅需原数据量25%即可训练出更准确的模型,降低75%数据需求。
  • 设计流式训练流程,在异构平台上快20%且精度不降,适合高效科研。

中子衍射结构确定任务计算成本高昂,通常需数小时至数天完成。近期研究表明,利用模拟中子散射数据训练机器学习模型可显著加速该过程。然而,模型需预测的结构参数越多,所需模拟数据量呈指数级增长,带来巨大计算挑战。为此,本文提出一种新型批处理主动学习(AL)策略,通过不确定性采样从概率分布中选取模型最不确定的样本进行标注。实验证明,该方法在仅使用约25%原始数据的情况下,仍能提升模型精度。随后,我们设计了一种基于该AL策略的高效流式训练工作流,并在两种异构平台上进行性能测试,结果表明,相比传统训练流程,该工作流可缩短约20%训练时间,且无精度损失。

原文摘要 · Abstract (English)

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with about 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers about 20% shorter training time without any loss of accuracy.

主动学习中子衍射结构预测流式训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。