利用咳嗽生理阶段设计自监督模型,提升声音分类准确性
CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

- 基于咳嗽的生理声学阶段构建正样本对进行自监督学习
- 在5个下游任务中表现优于传统随机裁剪方法,最高提升显著
- 适合呼吸疾病筛查与语音分析交叉研究者参考
本文提出CoughPhase-CLR,一种基于咳嗽生理阶段的自监督学习框架,用于鲁棒表征学习。不同于通用对比学习方法,该框架依据咳嗽的特定声学阶段构建正样本对。我们在约40小时公开咳嗽音频上预训练模型,并在五个下游任务中评估,包括新冠检测、慢性阻塞性肺病(COPD)状态分类和吸烟状态预测。结果表明,针对咳嗽的预训练始终优于标准随机裁剪方法。我们还对多种先进模型在COPD状态分类任务上进行了基准测试,最先进模型(在通用音频或呼吸音上预训练)的无加权平均召回率(UAR)仅为57%,远低于使用语音分析达到的84% UAR的现有最优性能。
原文摘要 · Abstract (English)
In this work, we introduce CoughPhase-CLR, a self-supervised learning framework designed to leverage the physiological phases of a cough for robust representation learning. Unlike generic contrastive frameworks, CoughPhase-CLR constructs positive pairs based on these specific acoustic phases. We pre-trained our model on approximately 40 hours of public cough audio and evaluated it across five downstream tasks, including COVID-19 detection, chronic obstructive pulmonary disease (COPD) state classification, and smoker status prediction. Our results demonstrate that cough-specific pre-training consistently outperforms standard random-cropping techniques when training on cough recordings. Additionally, we benchmarked a diverse set of state-of-the-art models on COPD state classification, highlighting the difficulty of this task. The best-performing models, pretrained on either general audio or respiratory sounds, achieved a UAR of 57\%, failing to outperform the state-of-the-art performance of 84\% UAR achieved using speech analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。