arXiv:2409.14312eess.AScs.SD2024-09被引 2

融合多种非语义语音特征,提升抑郁症检测准确率

Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection

  • 整合TRILLsson、x-vector等预训练模型提取的非语义特征
  • 在E-DAIC数据集上实现RMSE 5.51、MAE 4.48的顶尖表现
  • 方法简单有效,适合临床辅助诊断场景

本研究聚焦于从语音中识别抑郁症,关注非语义特征(NSFs)捕捉抑郁细微标志的潜力。尽管已有研究使用多种特征,但源自语音情感、说话人识别和副语言处理等预训练模型(如TRILLsson、x-vector、emoHuBERT)的非语义特征尚未被充分整合。本文证明,不同非语义特征具有互补性,可提升检测性能。为此,提出一种简单有效的融合框架FuSeR。实验表明,FuSeR优于单一特征模型及传统融合方法,在E-DAIC基准上达到RMSE 5.51、MAE 4.48,达到当前最佳水平,验证了其在抑郁症检测中的鲁棒性。

原文摘要 · Abstract (English)

In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection.

抑郁症检测语音分析特征融合深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。