用仿真数据训练的模型能更准识别ARPES谱图质量
A simulation-based training framework for machine-learning applications in ARPES
- 开发开源仿真工具aurelia生成大量虚拟ARPES数据
- 训练的CNN模型比人眼判断更准,且快速定位最佳测量区
- 适合需要高效分析ARPES数据的实验团队
近年来,角分辨光电子能谱(ARPES)在探测多维可观测量和生成多维数据集方面取得显著进展,但随之带来数据采集、处理与分析的新挑战。机器学习(ML)可大幅减轻实验人员负担,但缺乏用于训练的数据,尤其是深度学习所需数据,成为主要障碍。本文提出开源仿真工具aurelia,用于生成大规模合成ARPES谱图,以解决数据短缺问题。作为示范,我们训练卷积神经网络(CNN)评估ARPES谱图质量,该任务是实验初始样品对齐阶段的关键步骤。将仿真训练模型与真实实验数据对比后发现,该模型在谱图质量评估上优于人工判断,并能高精度快速定位最优测量区域。结果表明,仿真生成的ARPES谱图可有效替代真实数据用于机器学习模型训练。
原文摘要 · Abstract (English)
In recent years, angle-resolved photoemission spectroscopy (ARPES) has advanced significantly in its ability to probe more observables and simultaneously generate multi-dimensional datasets. These advances present new challenges in data acquisition, processing, and analysis. Machine learning (ML) models can drastically reduce the workload of experimentalists; however, the lack of training data for ML -- and in particular deep learning -- is a significant obstacle. In this work, we introduce an open-source synthetic ARPES spectra simulator - aurelia - for the purpose of generating the large datasets necessary to train ML models. As a demonstration, we train a convolutional neural network to evaluate ARPES spectra quality -- a critical task performed during the initial sample alignment phase of the experiment. We benchmark the simulation-trained model against actual experimental data and find that it can assess the spectra quality more accurately than human analysis, and swiftly identify the optimal measurement region with high precision. Thus, we establish that simulated ARPES spectra can be an effective proxy for experimental spectra in training ML models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。