融合波形与梅尔谱的神经网络提升呼吸音分类准确率
Waveform-Logmel Audio Neural Networks for Respiratory Sound Classification
- 同时输入原始波形和对数梅尔谱,用双向门控单元建模上下文特征
- 在SPRSound数据集上达到90.3%敏感度和93.6%总分
- 适合临床呼吸疾病辅助诊断,尤其适用于异常声音样本少的场景
使用电子听诊器进行听诊分析在呼吸系统疾病临床诊断中日益受到关注。近年来,神经网络已被用于辅助呼吸音分类并取得进展,但受限于异常呼吸音样本稀缺,仍具挑战。本文提出一种新架构——波形-对数梅尔谱音频神经网络(WLANN),同时以波形和对数梅尔谱为输入特征,并采用双向门控循环单元(Bi-GRU)对融合特征进行上下文建模。在SPRSound呼吸音数据集上的实验表明,所提框架能有效区分病理呼吸音类别,优于先前研究,敏感度达90.3%,总分为93.6%。本研究证实了WLANN在呼吸系统疾病诊断中的高有效性。
原文摘要 · Abstract (English)
Auscultatory analysis using an electronic stethoscope has attracted increasing attention in the clinical diagnosis of respiratory diseases. Recently, neural networks have been applied to assist in respiratory sound classification with achievements. However, it remains challenging due to the scarcity of abnormal respiratory sound. In this paper, we propose a novel architecture, namely Waveform-Logmel audio neural networks (WLANN), which uses both waveform and log-mel spectrogram as the input features and uses Bidirectional Gated Recurrent Units (Bi-GRU) to context model the fused features. Experimental results of our WLANN applied to SPRSound respiratory dataset show that the proposed framework can effectively distinguish pathological respiratory sound classes, outperforming the previous studies, with 90.3% in sensitivity and 93.6% in total score. Our study demonstrates the high effectiveness of the WLANN in the diagnosis of respiratory diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。