用自监督学习提升听诊音诊断能力,适合资源有限地区使用
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
- 通过自监督与对比学习融合多源数据训练
- 在新基准上异常检测准确率显著超越现有模型
- 开源模型代码,适合临床筛查与早期干预应用
准确高效的听诊诊断对早期疾病发现至关重要,尤其在临床专家稀缺的资源有限地区。传统听诊依赖医生经验,存在明显观察者差异;现有AI模型常因训练数据不具代表性而表现不佳。本研究提出AuscultaBase,一种基于自监督与对比学习的AI诊断框架,融合大规模多源数据以生成鲁棒特征表示,显著提升异常检测、疾病分类与活动识别性能。在新构建的基准AuscultaBench上全面评估显示,AuscultaBase在关键指标上持续优于现有最先进方法,展现出作为可扩展、低成本临床筛查工具的巨大潜力。代码与模型检查点已公开于https://github.com/applewpj/AuscultaBase。
原文摘要 · Abstract (English)
Accurate and efficient auscultation-based diagnostics are vital for early disease detection, especially in resource-limited settings where specialized clinical expertise is scarce. Traditional auscultation, which heavily depends on clinician experience, suffers from significant inter-observer variability, while existing AI models often falter due to the limitations of non-representative training data. In this study, we introduce AuscultaBase, a novel AI-driven diagnostic framework that harnesses self-supervised and contrastive learning techniques alongside large-scale, multi-source data integration to advance body sound analysis. By generating robust feature representations, AuscultaBase markedly enhances performance in abnormality detection, disease classification, and activity recognition tasks. Comprehensive evaluations on our newly established benchmark, AuscultaBench, demonstrate that AuscultaBase consistently outperforms state-of-the-art methods across key performance metrics, underscoring its potential as a scalable and cost-effective tool for clinical screening and early disease intervention. The code and model checkpoint has been released in https://github.com/applewpj/AuscultaBase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。