一个模型同时筛查9种声音相关疾病,准确率最高达0.97
Unified Acoustic Representations for Screening Neurological and Respiratory Pathologies from Voice
- 多任务学习框架共享声学骨干,跨病种知识迁移
- 在9种疾病上平均AUROC达0.78,阿尔茨海默症达0.97
- 无需原始音频,适合远程医疗与隐私敏感场景
基于语音的健康评估为大规模、非侵入式疾病筛查提供了新可能,但现有方法多针对单一病症,未能充分利用语音中丰富的多维度信息。我们提出MARVEL(多任务声学表征用于语音健康分析),一种注重隐私的多任务学习框架,仅使用提取的声学特征即可同时检测九类神经、呼吸及嗓音障碍,无需传输原始音频。其双分支架构采用专用编码器与任务特定头部,共享同一声学主干,实现有效的跨病种知识迁移。在大规模Bridge2AI-Voice v2.0数据集上,MARVEL整体AUROC达0.78,神经疾病表现优异(AUROC=0.89),尤其阿尔茨海默病/轻度认知障碍达0.97。该框架相比单模态基线提升5%-19%,在7项任务上超越当前自监督模型;相关性分析显示,模型学习的表征与临床已知声学特征具显著一致性,表明其内部表示符合医学认知模式。本研究证明单一统一模型可有效筛查多种疾病,为资源匮乏与远程医疗环境中的可部署语音诊断奠定基础。
原文摘要 · Abstract (English)
Voice-based health assessment offers unprecedented opportunities for scalable, non-invasive disease screening, yet existing approaches typically focus on single conditions and fail to leverage the rich, multi-faceted information embedded in speech. We present MARVEL (Multi-task Acoustic Representations for Voice-based Health Analysis), a privacy-conscious multitask learning framework that simultaneously detects nine distinct neurological, respiratory, and voice disorders using only derived acoustic features, eliminating the need for raw audio transmission. Our dual-branch architecture employs specialized encoders with task-specific heads sharing a common acoustic backbone, enabling effective cross-condition knowledge transfer. Evaluated on the large-scale Bridge2AI-Voice v2.0 dataset, MARVEL achieves an overall AUROC of 0.78, with exceptional performance on neurological disorders (AUROC = 0.89), particularly for Alzheimer's disease/mild cognitive impairment (AUROC = 0.97). Our framework consistently outperforms single-modal baselines by 5-19% and surpasses state-of-the-art self-supervised models on 7 of 9 tasks, while correlation analysis reveals that the learned representations exhibit meaningful similarities with established acoustic features, indicating that the model's internal representations are consistent with clinically recognized acoustic patterns. By demonstrating that a single unified model can effectively screen for diverse conditions, this work establishes a foundation for deployable voice-based diagnostics in resource-constrained and remote healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。