用生成模型增广医学音频数据,结果对分类效果提升有限。
Synthetic Data Augmentation for Medical Audio Classification: A Preliminary Evaluation
- 用VAE、GAN和扩散模型生成音频数据进行增广
- 基础模型F1为0.645,增广后多数情况未提升甚至下降
- 只有模型集成才小幅提升至0.664,适合后续优化研究
医学音频分类因信噪比低、特征细微、类内差异大,常受类别不平衡和训练数据少的制约。合成数据增广被视为缓解此问题的潜在策略,但以往研究方法不一,结果参差。本初步研究在中度不平衡数据集(73%:27%)上,使用基准深度卷积神经网络评估三种生成式增广策略(变分自编码器、生成对抗网络、扩散模型)的影响。无增广时基线模型F1得分为0.645。各增广策略均未带来性能提升,部分配置表现更差。仅模型集成方案实现小幅改善,F1达0.664。结果表明,对标准CNN分类器而言,合成增广在医学音频任务中未必有效。未来工作需关注任务特定数据特性、模型与增广的适配性及评估框架。
原文摘要 · Abstract (English)
Medical audio classification remains challenging due to low signal-to-noise ratios, subtle discriminative features, and substantial intra-class variability, often compounded by class imbalance and limited training data. Synthetic data augmentation has been proposed as a potential strategy to mitigate these constraints; however, prior studies report inconsistent methodological approaches and mixed empirical results. In this preliminary study, we explore the impact of synthetic augmentation on respiratory sound classification using a baseline deep convolutional neural network trained on a moderately imbalanced dataset (73%:27%). Three generative augmentation strategies (variational autoencoders, generative adversarial networks, and diffusion models) were assessed under controlled experimental conditions. The baseline model without augmentation achieved an F1-score of 0.645. Across individual augmentation strategies, performance gains were not observed, with several configurations demonstrating neutral or degraded classification performance. Only an ensemble of augmented models yielded a modest improvement in F1-score (0.664). These findings suggest that, for medical audio classification, synthetic augmentation may not consistently enhance performance when applied to a standard CNN classifier. Future work should focus on delineating task-specific data characteristics, model-augmentation compatibility, and evaluation frameworks necessary for synthetic augmentation to be effective in medical audio applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。