用自监督学习分析语音与面部动作,自动判断精神分裂症症状类型和严重程度。
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
- 基于VQ-VAE的多模态表征学习,从声道变量和面部动作单元提取通用特征。
- 多任务学习模型在分类任务上各项指标均优于此前方法,且首次实现严重程度评分。
- 适合精神健康领域研究者及临床辅助诊断系统开发者使用。
近年来,多模态精神分裂症评估系统受到关注。本文提出一种新框架,可区分精神分裂症主要症状类别并预测整体严重程度评分。我们构建了一个基于向量量化变分自编码器(VQ-VAE)的多模态表征学习(MRL)模型,从声道变量(TVs)和面部动作单元(FAUs)中生成任务无关的语音表征。这些表征被用于多任务学习(MTL)下游模型,以获得症状类别标签和总体严重程度分数。所提框架在多分类任务中各项评估指标(加权F1、AUC-ROC、加权准确率)均优于先前工作。此外,该模型首次实现了对精神分裂症严重程度的定量估计,这是以往方法未涉及的任务。
原文摘要 · Abstract (English)
Multimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variational Auto-Encoder (VQ-VAE) based Multimodal Representation Learning (MRL) model to produce task-agnostic speech representations from vocal Tract Variables (TVs) and Facial Action Units (FAUs). These representations are then used in a Multi-Task Learning (MTL) based downstream prediction model to obtain class labels and an overall severity score. The proposed framework outperforms the previous works on the multi-class classification task across all evaluation metrics (Weighted F1 score, AUC-ROC score, and Weighted Accuracy). Additionally, it estimates the schizophrenia severity score, a task not addressed by earlier approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。