用少量实测数据让仿真模型精准识别放射性同位素,突破模拟与现实的差距。
Sim-to-real supervised domain adaptation for radioisotope identification
- 用预训练+微调的迁移学习策略,在仿真与真实数据间架桥。
- 仅64个实测样本即达96%准确率,远超纯仿真或从零训练模型。
- 模型特征更易解释,适合辐射探测等数据稀缺场景。
机器学习有望提升伽马光谱法识别放射性同位素的速度与可靠性。然而,人工标注实验数据成本高昂;而仅用仿真数据训练则因模拟与真实测量间的域差异存在风险。本研究证明,监督域适应可显著提升识别模型性能,实现仿真与实验数据域之间的知识迁移。我们考虑两种场景:(1) 模拟到模拟的多标签比例估计,使用高纯锗探测器仿真数据;(2) 模拟到实验的多类单标签分类,使用便携式溴化镧(LaBr)和碘化钠(NaI)探测器的真实谱图。首先在仿真数据上预训练基于自定义Transformer的光谱分类器,再在仅64个标注实验谱图上微调,即在LaBr探测器的模拟到真实场景中达到96%测试准确率,大幅超越仅用仿真数据训练的基线模型(75%)和在相同64个样本上从零训练的模型(80%)。此外,域适应模型学习到的特征比纯实验基线模型更具人类可解释性。结果表明,监督域适应技术能有效弥合放射性同位素识别中的模拟到现实差距,即便在实验数据有限的真实场景下,也能构建高精度且可解释的分类器。
原文摘要 · Abstract (English)
Machine learning has the potential to improve the speed and reliability of radioisotope identification using gamma spectroscopy. However, meticulously labeling an experimental dataset for training is often prohibitively expensive, while training models purely on synthetic data is risky due to the domain gap between simulated and experimental measurements. In this research, we demonstrate that supervised domain adaptation can substantially improve the performance of radioisotope identification models by transferring knowledge between synthetic and experimental data domains. We consider two domain adaptation scenarios: (1) a simulation-to-simulation adaptation, where we perform multi-label proportion estimation using simulated high-purity germanium detectors, and (2) a simulation-to-experimental adaptation, where we perform multi-class, single-label classification using measured spectra from handheld lanthanum bromide (LaBr) and sodium iodide (NaI) detectors. We begin by pretraining a spectral classifier on synthetic data using a custom transformer-based neural network. After subsequent fine-tuning on just 64 labeled experimental spectra, we achieve a test accuracy of 96% in the sim-to-real scenario with a LaBr detector, far surpassing a synthetic-only baseline model (75%) and a model trained from scratch (80%) on the same 64 spectra. Furthermore, we demonstrate that domain-adapted models learn more human-interpretable features than experiment-only baseline models. Overall, our results highlight the potential for supervised domain adaptation techniques to bridge the sim-to-real gap in radioisotope identification, enabling the development of accurate and explainable classifiers even in real-world scenarios where access to experimental data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。