arXiv:2507.20408eess.SPcs.AI2025-07被引 5

用混合模型提升儿童肺部声音自动分类准确率

A Multi-Stage Hybrid CNN-Transformer Network for Automated Pediatric Lung Sound Classification

  • 融合卷积与注意力机制,处理儿童呼吸音的多阶段特征
  • 事件级分类准确率达0.9039(二分类)和0.8448(多分类)
  • 适合资源有限地区儿童呼吸疾病筛查,对低龄患儿有效

自动化肺部听诊分析对监测呼吸健康至关重要,尤其在医疗人员短缺地区。尽管成人呼吸音分类已广泛研究,但针对6岁以下儿童的研究仍不足。儿童肺部发育导致呼吸音声学特性显著变化,需专用分类方法。为此,我们提出一种多阶段混合CNN-Transformer框架,利用小波变换生成的谱图图像,从完整录音和单次呼吸中提取特征,进行儿童呼吸系统疾病分类。模型在事件级别二分类中取得0.9039的综合得分,多分类得分为0.8448,采用类别加权焦点损失缓解数据不平衡问题。在录音级别,三分类和多分类得分分别为0.720和0.571,分别优于此前最佳模型3.81%和5.94%。该方法为资源有限地区儿童呼吸疾病诊断提供了可行方案。

原文摘要 · Abstract (English)

Automated analysis of lung sound auscultation is essential for monitoring respiratory health, especially in regions facing a shortage of skilled healthcare workers. While respiratory sound classification has been widely studied in adults, its ap plication in pediatric populations, particularly in children aged <6 years, remains an underexplored area. The developmental changes in pediatric lungs considerably alter the acoustic proper ties of respiratory sounds, necessitating specialized classification approaches tailored to this age group. To address this, we propose a multistage hybrid CNN-Transformer framework that combines CNN-extracted features with an attention-based architecture to classify pediatric respiratory diseases using scalogram images from both full recordings and individual breath events. Our model achieved an overall score of 0.9039 in binary event classifi cation and 0.8448 in multiclass event classification by employing class-wise focal loss to address data imbalance. At the recording level, the model attained scores of 0.720 for ternary and 0.571 for multiclass classification. These scores outperform the previous best models by 3.81% and 5.94%, respectively. This approach offers a promising solution for scalable pediatric respiratory disease diagnosis, especially in resource-limited settings.

儿童医疗肺音分类深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。