分离疾病相关与无关特征,提升呼吸音分类鲁棒性。
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
- 双编码器分别提取疾病相关和无关特征,实现特征解耦。
- 在ICBHI数据集上表现优异,缓解设备与环境差异带来的域偏移。
- 适合医疗音频分析、跨设备呼吸音识别场景使用。
深度神经网络已用于音频频谱图进行呼吸音分类,但受限于数据稀缺,性能仍不理想。此外,由于呼吸音样本来自不同电子听诊器、患者群体和录音环境,模型训练中易引入领域差异。为此,我们提出一种改进的掩码自编码器(MAE)模型——解耦双编码器掩码自编码器(DDE-MAE),设计两个独立编码器分别捕捉疾病相关与无关信息,实现特征解耦以降低领域差异影响。该方法在ICBHI数据集上取得具有竞争力的性能。
原文摘要 · Abstract (English)
Deep neural networks have been applied to audio spectrograms for respiratory sound classification, but it remains challenging to achieve satisfactory performance due to the scarcity of available data. Moreover, domain mismatch may be introduced into the trained models as a result of the respiratory sound samples being collected from various electronic stethoscopes, patient demographics, and recording environments. To tackle this issue, we proposed a modified MaskedAutoencoder(MAE) model, named Disentangling Dual-Encoder MAE (DDE-MAE) for respiratory sound classification. Two independent encoders were designed to capture disease-related and disease-irrelevant information separately, achieving feature disentanglement to reduce the domain mismatch. Our method achieves a competitive performance on the ICBHI dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。