用语音中的元音特征,结合集成学习,更可靠地识别抑郁症状。
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
- 基于元音嵌入构建语音特征,融合语言语义信息
- 两种集成方法均达到顶尖水平,且抗数据分布偏差能力强
- 结果可解释,适合辅助临床诊断与筛查
本研究探讨了从语音中识别抑郁的可解释机器学习算法。基于抑郁症影响发音运动控制和元音生成的证据,采用预训练的元音嵌入,整合具有语义意义的语言单元。随后,通过集成学习将问题分解为由特定抑郁症状和严重程度表征的子任务。探索了两种方法:一种是自下而上的8模型方案,分别预测患者健康问卷-8项(PHQ-8)的单项得分;另一种是自上而下的专家混合(Mixture of Experts, MoE)方法,包含路由模块评估抑郁严重程度。两种方法性能均与当前最优基线相当,展现出强鲁棒性,且对数据集均值/中位数不敏感。讨论了系统可解释性的临床价值,强调其在辅助抑郁诊断与筛查中的潜力。
原文摘要 · Abstract (English)
This study investigates explainable machine learning algorithms for identifying depression from speech. Grounded in evidence from speech production that depression affects motor control and vowel generation, pre-trained vowel-based embeddings, that integrate semantically meaningful linguistic units, are used. Following that, an ensemble learning approach decomposes the problem into constituent parts characterized by specific depression symptoms and severity levels. Two methods are explored: a "bottom-up" approach with 8 models predicting individual Patient Health Questionnaire-8 (PHQ-8) item scores, and a "top-down" approach using a Mixture of Experts (MoE) with a router module for assessing depression severity. Both methods depict performance comparable to state-of-the-art baselines, demonstrating robustness and reduced susceptibility to dataset mean/median values. System explainability benefits are discussed highlighting their potential to assist clinicians in depression diagnosis and screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。