用语音和文字联合分析,提前精准识别抑郁和创伤后应激障碍。
Innovative Framework for Early Estimation of Mental Disorder Scores to Enable Timely Interventions
- 融合文本与音频特征,用LSTM和BiLSTM双模建模。
- 抑郁分类准确率92%,PTSD达93%,优于单一模态方法。
- 适合临床早期筛查,助力心理疾病及时干预。
抑郁症和创伤后应激障碍(PTSD)严重影响个体整体福祉,因此早期检测与精确诊断对及时临床干预至关重要。本文提出一种先进的多模态深度学习系统,用于自动分类PTSD与抑郁。该方法利用临床访谈数据集中的文本与音频数据,结合LSTM(长短期记忆)与BiLSTM(双向长短期记忆)架构,提取双模态特征:文本特征关注语义与语法成分,音频特征捕捉语调、节奏与音高等声学特性。多模态融合增强了模型识别心理健康细微模式的能力。在测试集上,该方法对抑郁的分类准确率达92%,对PTSD达93%,显著优于传统单模态方法,展现出更高的准确性和鲁棒性。
原文摘要 · Abstract (English)
Individual's general well-being is greatly impacted by mental health conditions including depression and Post-Traumatic Stress Disorder (PTSD), underscoring the importance of early detection and precise diagnosis in order to facilitate prompt clinical intervention. An advanced multimodal deep learning system for the automated classification of PTSD and depression is presented in this paper. Utilizing textual and audio data from clinical interview datasets, the method combines features taken from both modalities by combining the architectures of LSTM (Long Short Term Memory) and BiLSTM (Bidirectional Long Short-Term Memory).Although text features focus on speech's semantic and grammatical components; audio features capture vocal traits including rhythm, tone, and pitch. This combination of modalities enhances the model's capacity to identify minute patterns connected to mental health conditions. Using test datasets, the proposed method achieves classification accuracies of 92% for depression and 93% for PTSD, outperforming traditional unimodal approaches and demonstrating its accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。