用咳嗽声+语音大模型,实现低成本高精度结核病筛查。
Deep Learning for Tuberculosis Screening in a High-burden Setting using Cough Analysis and Speech Foundation Models
- 用预训练语音模型微调咳嗽音频,识别结核病特征。
- 多模态融合达92.1%准确率,90.3%敏感性,符合世卫标准。
- 模型抗噪声、设备差异等干扰,适合资源匮乏地区使用。
人工智能可识别咳嗽声中的疾病特征,为高负担、资源有限地区提供可扩展、低成本的结核病(TB)筛查方案。以往研究受限于小样本、非结核症状患者代表性不足,以及在受控环境中采集数据。本研究在赞比亚两家医院招募512名参与者,分为三组:细菌学确诊结核病(TB+)、有呼吸道症状的其他疾病患者(OR)、健康对照(HC)。获得500名参与者的可用咳嗽录音及临床数据。基于预训练语音基础模型的深度学习分类器,在3秒音频片段上微调以预测诊断类别。最优模型在区分TB咳嗽与所有其他人群(TB+/Rest)时达到85.2% AUROC,TB+ vs. OR患者时为80.1%。融合人口统计与临床特征后,性能提升至92.1%(TB+/Rest)和84.2%(TB+/OR)。在概率阈值0.38下,多模态模型对TB+/Rest的敏感性达90.3%,特异性73.1%,满足世卫组织结核病筛查目标产品要求。对抗测试与分层分析显示,模型对背景噪声、录音时间、设备差异等混杂因素具有鲁棒性。结果证明,基于咳嗽的AI在真实低资源环境下具备可行性。
原文摘要 · Abstract (English)
Artificial intelligence (AI) systems can detect disease-related acoustic patterns in cough sounds, offering a scalable and cost-effective approach to tuberculosis (TB) screening in high-burden, resource-limited settings. Previous studies have been limited by small datasets, under-representation of symptomatic non-TB patients, and recordings collected in controlled environments. In this study, we enrolled 512 participants at two hospitals in Zambia, categorised into three groups: bacteriologically confirmed TB (TB+), symptomatic patients with other respiratory diseases (OR), and healthy controls (HC). Usable cough recordings with demographic and clinical data were obtained from 500 participants. Deep learning classifiers based on pre-trained speech foundation models were fine-tuned on cough recordings to predict diagnostic categories. The best-performing model, trained on 3-second audio clips, achieved an AUROC of 85.2% for distinguishing TB coughs from all other participants (TB+/Rest) and 80.1% for TB+ versus symptomatic OR participants (TB+/OR). Incorporating demographic and clinical features improved performance to 92.1% for TB+/Rest and 84.2% for TB+/OR. At a probability threshold of 0.38, the multimodal model reached 90.3% sensitivity and 73.1% specificity for TB+/Rest, meeting WHO target product profile benchmarks for TB screening. Adversarial testing and stratified analyses shows that the model was robust to confounding factors including background noise, recording time, and device variability. These results demonstrate the feasibility of cough-based AI for TB screening in real-world, low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。