用呼吸声自动识别肺病,还能解释判断依据。
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
- 融合深度学习与人工声学特征,双路编码后晚期融合。
- 在哮喘数据集上达91.21%准确率,优于所有消融实验。
- 支持图像级、时间级、特征级可解释性,适合医疗场景。
呼吸系统疾病仍是全球重大健康挑战,传统听诊受主观性、环境噪声和医生间差异限制。本研究提出一种可解释的多模态深度学习框架,基于呼吸音频信号实现肺病自动检测。系统整合两种互补表征:基于CNN-BiLSTM注意力架构的谱时序编码器,以及捕捉生理意义声学特征(如MFCCs、频谱质心、频谱带宽、过零率)的手工特征编码器。两条分支通过晚期融合,结合数据驱动学习与领域先验声学线索。模型在Asthma Detection Dataset Version 2上训练与评估,经重采样、归一化、降噪、数据增强及患者级分层划分等严格预处理。结果表现出强泛化能力:准确率91.21%,宏平均F1-score为0.899,宏平均ROC-AUC达0.9866,优于所有消融变体。消融实验验证了时序建模、注意力机制和多模态融合的重要性。框架引入Grad-CAM、Integrated Gradients与SHAP,生成与已知声学生物标志物对齐的谱图、时间与特征级可解释性,提升临床透明度。研究结果表明该框架在远程医疗、床旁诊断与真实世界呼吸筛查中具有应用潜力。
原文摘要 · Abstract (English)
Respiratory diseases remain major global health challenges, and traditional auscultation is often limited by subjectivity, environmental noise, and inter-clinician variability. This study presents an explainable multimodal deep learning framework for automatic lung-disease detection using respiratory audio signals. The proposed system integrates two complementary representations: a spectral-temporal encoder based on a CNN-BiLSTM Attention architecture, and a handcrafted acoustic-feature encoder capturing physiologically meaningful descriptors such as MFCCs, spectral centroid, spectral bandwidth, and zero-crossing rate. These branches are combined through late-stage fusion to leverage both data-driven learning and domain-informed acoustic cues. The model is trained and evaluated on the Asthma Detection Dataset Version 2 using rigorous preprocessing, including resampling, normalization, noise filtering, data augmentation, and patient-level stratified partitioning. The study achieved strong generalization with 91.21% accuracy, 0.899 macro F1-score, and 0.9866 macro ROC-AUC, outperforming all ablated variants. An ablation study confirms the importance of temporal modeling, attention mechanisms, and multimodal fusion. The framework incorporates Grad-CAM, Integrated Gradients, and SHAP, generating interpretable spectral, temporal, and feature-level explanations aligned with known acoustic biomarkers to build clinical transparency. The findings demonstrate the framework's potential for telemedicine, point-of-care diagnostics, and real-world respiratory screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。