arXiv:2606.21411cs.SD2026-06

利用咳嗽生理阶段设计自监督模型,提升声音分类准确性

CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

论文配图:CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification
图 1 · 摘自论文原文
  • 基于咳嗽的生理声学阶段构建正样本对进行自监督学习
  • 在5个下游任务中表现优于传统随机裁剪方法,最高提升显著
  • 适合呼吸疾病筛查与语音分析交叉研究者参考

本文提出CoughPhase-CLR,一种基于咳嗽生理阶段的自监督学习框架,用于鲁棒表征学习。不同于通用对比学习方法,该框架依据咳嗽的特定声学阶段构建正样本对。我们在约40小时公开咳嗽音频上预训练模型,并在五个下游任务中评估,包括新冠检测、慢性阻塞性肺病(COPD)状态分类和吸烟状态预测。结果表明,针对咳嗽的预训练始终优于标准随机裁剪方法。我们还对多种先进模型在COPD状态分类任务上进行了基准测试,最先进模型(在通用音频或呼吸音上预训练)的无加权平均召回率(UAR)仅为57%,远低于使用语音分析达到的84% UAR的现有最优性能。

原文摘要 · Abstract (English)

In this work, we introduce CoughPhase-CLR, a self-supervised learning framework designed to leverage the physiological phases of a cough for robust representation learning. Unlike generic contrastive frameworks, CoughPhase-CLR constructs positive pairs based on these specific acoustic phases. We pre-trained our model on approximately 40 hours of public cough audio and evaluated it across five downstream tasks, including COVID-19 detection, chronic obstructive pulmonary disease (COPD) state classification, and smoker status prediction. Our results demonstrate that cough-specific pre-training consistently outperforms standard random-cropping techniques when training on cough recordings. Additionally, we benchmarked a diverse set of state-of-the-art models on COPD state classification, highlighting the difficulty of this task. The best-performing models, pretrained on either general audio or respiratory sounds, achieved a UAR of 57\%, failing to outperform the state-of-the-art performance of 84\% UAR achieved using speech analysis.

声音分类自监督学习呼吸疾病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。