用自编码器生成2D光谱数据,解决标注样本少的问题。
Synthetic generation of 2D data records based on Autoencoders
- 基于自编码器框架生成合成2D光谱数据
- 加入合成数据后分类准确率显著提升
- 适合标注数据稀缺的光谱分析场景
气相色谱-离子迁移谱(GC-IMS)是一种双分离分析技术,广泛用于气体样品组分识别,通过分离并分析组分到达时间获得二维谱图。此类数据虽信息丰富,但因标注数据集有限,给数据驱动分析带来挑战。本研究提出一种基于深度学习自编码器的新型2D谱图合成方法。该方法虽应用于GC-IMS数据,但可推广至任意二维光谱测量场景,当在标注数据集上进行组分分类时,引入合成记录显著提升了分类性能,验证了其在克服机器学习中数据集限制方面的潜力。
原文摘要 · Abstract (English)
Gas Chromatography coupled with Ion Mobility Spectrometry (GC-IMS) is a dual-separation analytical technique widely used for identifying components in gaseous samples by separating and analysing the arrival times of their constituent species. Data generated by GC-IMS is typically represented as two-dimensional spectra, providing rich information but posing challenges for data-driven analysis due to limited labelled datasets. This study introduces a novel method for generating synthetic 2D spectra using a deep learning framework based on Autoencoders. Although applied here to GC-IMS data, the approach is broadly applicable to any two-dimensional spectral measurements where labelled data are scarce. While performing component classification over a labelled dataset of GC-IMS records, the addition of synthesized records significantly has improved the classification performance, demonstrating the method's potential for overcoming dataset limitations in machine learning frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。