对比机器学习与统计方法在伽马能谱自动识别中的表现
Comparative study of machine learning and statistical methods for automatic identification and quantification in γ-ray spectrometry
- 构建开放基准测试平台,包含模拟数据与多种分析代码
- 统计方法在所有场景下识别准确率均优于机器学习
- 当谱线模型不准时,统计法性能下降,机器学习更鲁棒
过去十年中,众多数值方法被提出用于伽马能谱的自动识别与量化。然而,缺乏统一的基准(如数据集、代码、评估指标)使得方法比较困难。为此,我们提出一个开源基准,包含多种伽马能谱设置的模拟数据集、不同分析方法的代码及评估指标。在三个场景下进行比较:(1)谱线已知;(2)因康普顿散射和衰减导致谱线形变;(3)谱线因温度变化发生偏移。每个场景使用含9种核素和实验本底的20万条模拟谱数据,多核素共存。结果显示,统计方法在所有场景下识别性能均优于端到端机器学习方法。但若谱线模型不准确,统计方法性能显著下降。因此,全谱统计法在谱线已知或建模良好时最优,而机器学习适用于测量条件不确定的情况。量化任务中,统计方法可提供精确计数估计,机器学习结果则较差。
原文摘要 · Abstract (English)
During the last decade, a large number of different numerical methods have been proposed to tackle the automatic identification and quantification in γ-ray spectrometry. However, the lack of common benchmarks, including datasets, code and comparison metrics, makes their evaluation and comparison hard. In that context, we propose an open-source benchmark that comprises simulated datasets of various γ-spectrometry settings, codes of different analysis approaches and evaluation metrics. This allows us to compare the state-of-the-art end-to-end machine learning with a statistical unmixing approach using the full spectrum. Three scenarios have been investigated: (1) spectral signatures are assumed to be known; (2) spectral signatures are deformed due to physical phenomena such as Compton scattering and attenuation; and (3) spectral signatures are shifted (e.g., due to temperature variation). A large dataset of 200000 simulated spectra containing nine radionuclides with an experimental natural background is used for each scenario with multiple radionuclides present in the spectrum. Regarding identification performance, the statistical approach consistently outperforms the machine learning approaches across all three scenarios for all comparison metrics. However, the performance of the statistical approach can be significantly impacted when spectral signatures are not modeled correctly. Consequently, the full-spectrum statistical approach is most effective with known or well-modeled spectral signatures, while end-to-end machine learning is a good alternative when measurement conditions are uncertain for radionuclide identification. Concerning the quantification task, the statistical approach provides accurate estimates of radionuclide counting, while the machine learning methods deliver less satisfactory results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。