轻量级框架提升呼吸音分类准确率,适配移动端部署
CycleGuardian: A Framework for Automatic RespiratorySound classification Based on Improved Deep clustering and Contrastive Learning
- 融合改进的深度聚类与对比学习,增强异常音特征捕捉
- 在ICBHI2017数据集上达到63.26%综合得分,领先当前模型近7%
- 模型仅38M,可直接部署于安卓设备,支持移动听诊应用
听诊在呼吸系统疾病早期诊断中至关重要。尽管新冠后深度学习方法兴起,但受限于小规模数据集,性能提升受阻。正常与异常呼吸音共存时,正常成分与噪声干扰区分困难;不同异常音具有相似异常特征,难以辨别。此外,现有模型参数量过大,难在资源受限的移动平台部署。为此,本文设计轻量网络CycleGuardian,提出基于改进深度聚类与对比学习的框架。首先生成混合频谱图以增强特征多样性,并分组频谱图以捕捉间歇性异常音。其次,整合深度聚类模块与相似性约束聚类组件,提升异常特征捕获能力;引入对比学习模块与组内混合策略,增强异常特征辨识度。多目标优化提升训练整体表现。实验采用ICBHI2017数据集,按官方划分方式且不使用预训练权重,本方法达Sp: 82.06%,Se: 44.47%,Score: 63.26%,模型大小仅38M,相较当前最优模型提升近7%,达到当前最佳性能。同时,已成功部署于安卓设备,实现完整的智能呼吸音听诊系统。
原文摘要 · Abstract (English)
Auscultation plays a pivotal role in early respiratory and pulmonary disease diagnosis. Despite the emergence of deep learning-based methods for automatic respiratory sound classification post-Covid-19, limited datasets impede performance enhancement. Distinguishing between normal and abnormal respiratory sounds poses challenges due to the coexistence of normal respiratory components and noise components in both types. Moreover, different abnormal respiratory sounds exhibit similar anomalous features, hindering their differentiation. Besides, existing state-of-the-art models suffer from excessive parameter size, impeding deployment on resource-constrained mobile platforms. To address these issues, we design a lightweight network CycleGuardian and propose a framework based on an improved deep clustering and contrastive learning. We first generate a hybrid spectrogram for feature diversity and grouping spectrograms to facilitating intermittent abnormal sound capture.Then, CycleGuardian integrates a deep clustering module with a similarity-constrained clustering component to improve the ability to capture abnormal features and a contrastive learning module with group mixing for enhanced abnormal feature discernment. Multi-objective optimization enhances overall performance during training. In experiments we use the ICBHI2017 dataset, following the official split method and without any pre-trained weights, our method achieves Sp: 82.06 $\%$, Se: 44.47$\%$, and Score: 63.26$\%$ with a network model size of 38M, comparing to the current model, our method leads by nearly 7$\%$, achieving the current best performances. Additionally, we deploy the network on Android devices, showcasing a comprehensive intelligent respiratory sound auscultation system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。