用耳蜗图+视觉变换器识别呼吸杂音,提升诊断效率。
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers
- 将耳蜗图输入视觉变换器,捕捉声音的时频特征。
- 在ICBHI数据集上分类准确率超越传统CNN方法。
- 适合医疗智能诊断、呼吸疾病自动筛查场景。
早期识别呼吸异常对改善肺部健康、降低全球死亡率至关重要。呼吸音分析在评估呼吸系统状态和识别异常中发挥关键作用。本研究首次将耳蜗图作为输入,结合视觉变换器(ViT)进行杂音分类,探索该组合在该任务中的表现。尽管ViT已在音频分类中通过自注意力机制处理频谱图块取得良好效果,但本文进一步采用耳蜗图,以捕捉杂音特有的时频特征。所提方法在ICBHI数据集上进行评估,并与使用频谱图、梅尔频率倒谱系数、常Q变换及耳蜗图作为输入的其他先进CNN方法进行比较。结果表明,耳蜗图与ViT结合可实现更优分类性能,凸显了ViT在可靠呼吸音分类中的潜力。本研究推动了自动化智能技术的发展,旨在显著提升呼吸疾病检测的速度与效果,满足医学领域的迫切需求。
原文摘要 · Abstract (English)
Early identification of respiratory irregularities is critical for improving lung health and reducing global mortality rates. The analysis of respiratory sounds plays a significant role in characterizing the respiratory system's condition and identifying abnormalities. The main contribution of this study is to investigate the performance when the input data, represented by cochleogram, is used to feed the Vision Transformer architecture, since this input classifier combination is the first time it has been applied to adventitious sound classification to our knowledge. Although ViT has shown promising results in audio classification tasks by applying self attention to spectrogram patches, we extend this approach by applying the cochleogram, which captures specific spectro-temporal features of adventitious sounds. The proposed methodology is evaluated on the ICBHI dataset. We compare the classification performance of ViT with other state of the art CNN approaches using spectrogram, Mel frequency cepstral coefficients, constant Q transform, and cochleogram as input data. Our results confirm the superior classification performance combining cochleogram and ViT, highlighting the potential of ViT for reliable respiratory sound classification. This study contributes to the ongoing efforts in developing automatic intelligent techniques with the aim to significantly augment the speed and effectiveness of respiratory disease detection, thereby addressing a critical need in the medical field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。