用生成模型扩充小样本脑瘤光谱数据,提升分类准确率。
Generative Augmentation of Raman Spectra for Glioma Classification

- 构建条件变分自编码器生成特定类别的合成拉曼光谱。
- 合成数据加入真实数据后,多模型分类性能均提升。
- 适合小样本医学光谱分析,尤其数据稀缺的癌症诊断场景。
基于拉曼光谱的生物医学诊断中,机器学习面临数据集规模不足的挑战,尤其在脑胶质瘤分析中,数据量小且存在采集差异。本文基于58例肿瘤样本的活检光谱,研究深度生成式数据增强在小样本下的有效性。针对数据量少和类别不平衡问题,提出一种条件变分自编码器(β-CVAE),可生成类别相关的合成光谱。在严格患者隔离的交叉验证下,分别测试了仅用合成数据训练(TS/TR)和合成+真实数据联合训练(TSR/TR)的性能。结果显示,仅使用合成数据训练的模型性能低于真实数据训练模型,表明合成与真实数据分布间存在显著域偏移。但将合成数据作为真实数据的补充后,各类分类模型性能均有提升。此外,提出基于重构误差的分类方法(CbR),通过不同类别条件下的重构表现进行预测。结果表明,即使患者样本有限,生成模型仍能捕捉足够结构,为下游分类器提供有效正则化。该策略可提升拉曼光谱在小样本医学应用中的模型鲁棒性。
原文摘要 · Abstract (English)
Access to sufficiently large biomedical datasets remains a major obstacle for machine learning in Raman spectroscopy-based diagnostics. In particular, for glioma analysis, datasets are typically small and heterogeneous, affected by acquisition-specific variability. This work investigates the utility of deep generative augmentation in such a small-cohort setting. We analyze glioma biopsy spectra acquired from 58 tumor samples and consider both binary IDH-status classification and 6-class methylation subtype classification problems. To address the limited size and imbalance of the dataset, we develop a conditional variational autoencoder ($β$-CVAE) capable of generating class-conditioned synthetic Raman spectra. The generated data are evaluated in Train-on-Synthetic, Test-on-Real (TS/TR) and Train-on-Synthetic+Real, Test-on-Real (TSR/TR) settings under a strict patient-isolated cross-validation protocol. Models trained exclusively on synthetic data underperform models trained on real spectra, indicating a substantial domain gap between synthetic and real distributions. However, augmenting the real training data with synthetic spectra consistently improves classification performance across multiple models. These findings indicate that, even with a limited number of independent patient samples, generative models can capture sufficient structure to provide useful regularization for downstream classifiers. We also investigate a reconstruction-based inference strategy, termed Classification by Reconstruction (CbR), in which class prediction is based on reconstruction error under different class conditions. Overall, the results support the use of deep generative augmentation as a practical strategy for improving machine learning robustness in Raman spectroscopy applications characterized by limited biomedical datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。