arXiv:2411.06033eess.AS2024-11被引 6

融合语音发音与自监督特征,提升精神分裂症严重程度评估精度

Speech-Based Estimation of Schizophrenia Severity Using Feature Fusion

  • 用自监督模型提取语音发音特征,结合多类语音表示进行融合
  • 相比此前视听融合模型,误差降低9.18%(MAE)和9.36%(RMSE)
  • 适合精神健康智能评估、临床辅助诊断等场景使用

近年来,基于语音的精神分裂症谱系评估受到广泛关注。本研究提出一种深度学习框架,通过融合发音特征与来自预训练音频模型的多种自监督语音特征,实现对精神分裂症严重程度的估计。我们还设计了一种基于自编码器的自监督表征学习方法,从语音中提取紧凑的发音嵌入。性能最优的融合模型采用多头注意力机制,在严重程度预测任务上相比此前结合语音与视频输入的模型,均方误差(MAE)降低9.18%,均方根误差(RMSE)降低9.36%。

原文摘要 · Abstract (English)

Speech-based assessment of the schizophrenia spectrum has been widely researched over in the recent past. In this study, we develop a deep learning framework to estimate schizophrenia severity scores from speech using a feature fusion approach that fuses articulatory features with different self-supervised speech features extracted from pre-trained audio models. We also propose an auto-encoder-based self-supervised representation learning framework to extract compact articulatory embeddings from speech. Our top-performing speech-based fusion model with Multi-Head Attention (MHA) reduces Mean Absolute Error (MAE) by 9.18% and Root Mean Squared Error (RMSE) by 9.36% for schizophrenia severity estimation when compared with the previous models that combined speech and video inputs.

精神分裂症语音分析自监督学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。