arXiv:2508.03764cs.SDcs.AI2025-08中稿 · ISWC被引 1

用自监督学习提升咳嗽音频的通用表征,解决数据少时诊断不准问题。

CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning

  • 通过掩码数据建模实现自监督训练,无需人工标注即可学习咳嗽声特征。
  • 在3个咳嗽分类任务中表现媲美甚至超越现有监督模型。
  • 适合数据稀缺场景下的呼吸疾病早期诊断研究者使用。

医生在诊断过程中常通过听诊呼吸道声音来评估患者气道状况。近年来,基于人工智能的呼吸音诊断系统在呼吸系统疾病检测中展现出良好效果,推动了早期、便捷的诊断进展。然而,标签和数据稀缺仍是主要挑战,尤其在新冠以外的疾病上,限制了诊断性能与可靠评估。本文提出CoughViT,一种用于学习通用咳嗽声表征的新型预训练框架,以提升小样本任务下的诊断性能。为应对标签稀缺问题,采用掩码数据建模方法,以自监督方式训练特征编码器。我们在三个具有临床意义的咳嗽分类任务上评估该方法,结果表明,其表征在下游任务中表现达到或超过当前最优的监督音频表征。

原文摘要 · Abstract (English)

Physicians routinely assess respiratory sounds during the diagnostic process, providing insight into the condition of a patient's airways. In recent years, AI-based diagnostic systems operating on respiratory sounds, have demonstrated success in respiratory disease detection. These systems represent a crucial advancement in early and accessible diagnosis which is essential for timely treatment. However, label and data scarcity remain key challenges, especially for conditions beyond COVID-19, limiting diagnostic performance and reliable evaluation. In this paper, we propose CoughViT, a novel pre-training framework for learning general-purpose cough sound representations, to enhance diagnostic performance in tasks with limited data. To address label scarcity, we employ masked data modelling to train a feature encoder in a self-supervised learning manner. We evaluate our approach against other pre-training strategies on three diagnostically important cough classification tasks. Experimental results show that our representations match or exceed current state-of-the-art supervised audio representations in enhancing performance on downstream tasks.

自监督学习语音表征呼吸疾病诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。