用几何优化提升呼吸音分类敏感度,效果优于现有方法。
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
- 引入SAM优化音频频谱变换器的损失曲面几何,防止过拟合。
- 在ICBHI 2017上达到68.10%准确率和68.31%敏感度,创纪录。
- 适合医疗听诊场景,尤其关注高敏感度筛查需求。
呼吸音分类受限于ICBHI 2017等基准数据集规模小、噪声高、类别严重不平衡。尽管基于Transformer的模型具备强大特征提取能力,但在有限医学数据上训练时易过拟合,常收敛至损失曲面的尖锐极小值。为此,本文提出一种框架,通过Sharpness-Aware Minimization(SAM)优化音频频谱变换器(AST)的损失曲面几何,引导模型向更平坦的极小值收敛,从而提升泛化能力。同时采用加权采样策略有效应对类别不平衡问题。实验表明,该方法在ICBHI 2017数据集上取得68.10%的准确率,超越现有CNN与混合基线;更重要的是,敏感度达68.31%,显著提升临床筛查可靠性。t-SNE与注意力图分析进一步验证了模型学习到的是鲁棒且具区分性的特征,而非记忆背景噪声。
原文摘要 · Abstract (English)
Respiratory sound classification is hindered by the limited size, high noise levels, and severe class imbalance of benchmark datasets like ICBHI 2017. While Transformer-based models offer powerful feature extraction capabilities, they are prone to overfitting and often converge to sharp minima in the loss landscape when trained on such constrained medical data. To address this, we introduce a framework that enhances the Audio Spectrogram Transformer (AST) using Sharpness-Aware Minimization (SAM). Instead of merely minimizing the training loss, our approach optimizes the geometry of the loss surface, guiding the model toward flatter minima that generalize better to unseen patients. We also implement a weighted sampling strategy to handle class imbalance effectively. Our method achieves a state-of-the-art score of 68.10% on the ICBHI 2017 dataset, outperforming existing CNN and hybrid baselines. More importantly, it reaches a sensitivity of 68.31%, a crucial improvement for reliable clinical screening. Further analysis using t-SNE and attention maps confirms that the model learns robust, discriminative features rather than memorizing background noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。