arXiv:2412.19072eess.AScs.CL2024-12被引 5

用语音和文本双模型实现跨用户抑郁筛查,效果稳定可靠。

Robust Speech and Natural Language Processing Models for Depression Screening

  • 基于声学与自然语言的双模型,均采用迁移学习。
  • 在无重叠说话人数据上,二分类AUC均超0.80。
  • 对说话人和会话变量具有强鲁棒性,适合实际部署。

抑郁症是全球性的健康问题,亟需提升患者筛查覆盖率。语音技术为远程筛查提供了优势,但必须在不同患者间保持稳健表现。本文提出了两种深度学习模型:一种基于声学特征,另一种基于自然语言处理,均采用迁移学习。使用包含11,000名独特用户的抑郁标注语料库数据,该语料库通过人机对话语音交互构建。在二分类抑郁检测任务中,两个模型在未见数据上(无说话人重叠)均达到AUC≥0.80。进一步分析表明,模型在说话人和会话变量上的表现总体稳健。结论表明,该方法具备实现通用自动化抑郁筛查的潜力。

原文摘要 · Abstract (English)

Depression is a global health concern with a critical need for increased patient screening. Speech technology offers advantages for remote screening but must perform robustly across patients. We have described two deep learning models developed for this purpose. One model is based on acoustics; the other is based on natural language processing. Both models employ transfer learning. Data from a depression-labeled corpus in which 11,000 unique users interacted with a human-machine application using conversational speech is used. Results on binary depression classification have shown that both models perform at or above AUC=0.80 on unseen data with no speaker overlap. Performance is further analyzed as a function of test subset characteristics, finding that the models are generally robust over speaker and session variables. We conclude that models based on these approaches offer promise for generalized automated depression screening.

抑郁筛查语音分析NLP迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。