用深度学习从语音中提取抑郁焦虑生物标记,效果优于传统方法。
Voice Biomarkers for Depression and Anxiety

- 直接处理原始语音,自动学习内容无关的生物标记特征。
- 在约5000人测试集上达到71%的敏感度与特异度。
- 适合心理健康评估、临床辅助诊断等场景研究者使用。
当前基于语音检测抑郁和焦虑的方法主要依赖手工设计的副语言特征及声学描述符,这些特征来自语音信号的时间域与频域表示。将深度学习直接应用于原始语音信号,有望获得更具预测力的生物标记表征。然而,这类方法通常需要大量精心标注的数据,以学习鲁棒且具有临床意义的生物标记表示。本文介绍我们基于一个大规模专有数据集开发的深度学习模型,该数据集包含约6.5万条语音片段,来自超过2.3万名符合美国人口统计特征的受试者。我们展示了所采用的技术并分析其对模型性能的影响。结果表明,所提模型可提取与内容无关的生物标记信息,结合从音频中提取的词汇特征后,在实际应用环境中表现出更优的预测能力。模型在约5000名独立受试者上进行评估,达到71%的敏感度与特异度。为促进语音心理健康评估的研究,我们已将本文最佳模型发布至HuggingFace。
原文摘要 · Abstract (English)
Current approaches to detecting depression and anxiety from speech primarily rely on machine learning techniques that utilize hand-engineered paralinguistic features and related acoustic descriptors derived from time- and frequency-domain representations of speech signals. Applying deep learning methods directly to raw speech signals has the potential to produce biomarker representations with substantially greater predictive power. However, these approaches typically require large volumes of carefully annotated data to learn robust and clinically meaningful representations of the underlying biomarkers. In this paper, we describe our efforts toward developing a deep learning model trained on a large-scale proprietary dataset comprising ~65,000 utterances collected from more than 23,000 subjects representative of relevant United States demographics. We present the techniques employed and analyze their impact on model performance. Our results demonstrate that the proposed models can extract content-agnostic biomarker information, which, when combined with lexical features extracted from audio, yields improved predictive performance in production settings. Our models are evaluated on ~5000 unique subjects and achieve performance of 71% in terms of sensitivity and specificity. To foster further research in mental health assessment from speech, we release the best-performing model described in this paper on HuggingFace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。