用多模态嵌入技术提升抑郁症和创伤后应激障碍的客观评估准确率。
Leveraging Embedding Techniques in Multimodal Machine Learning for Mental Illness Assessment
- 按语句切片处理文本音频视频,增强模态信息提取效果。
- 决策层融合结合大模型预测,抑郁与PTSD检测准确率达94.8%和96.2%。
- 适合心理健康筛查、远程诊疗等需要可扩展评估工具的场景。
全球精神疾病(如抑郁、创伤后应激障碍)发病率持续上升,亟需客观且可扩展的诊断工具。传统临床评估存在可及性差、主观性强等问题。本文探索多模态机器学习在该领域的潜力,利用文本、音频、视频数据的互补信息。研究系统分析了多种数据预处理技术,包括创新的分段与基于语句的格式化策略;评估了各模态的先进嵌入模型,并采用卷积神经网络(CNN)与双向长短期记忆网络(BiLSTM)进行特征提取。比较了数据级、特征级与决策级融合方法,提出将大语言模型(LLM)预测结果融入决策流程的新方案。同时考察了以支持向量机替代多层感知机对分类性能的影响。拓展至使用PHQ-8与PCL-C评分的严重程度预测及共病多分类任务。结果显示,基于语句的分段显著提升性能,尤其在文本与音频模态中表现突出;决策级融合结合LLM预测取得最高精度,抑郁检测平衡准确率达94.8%,PTSD检测达96.2%。结合CNN-BiLSTM架构与语句级分段,再融入外部大模型,为精神健康状况的检测与评估提供了强大而细致的方法。研究证实多模态机器学习在开发更精准、可及、个性化的心理健康工具方面具有巨大潜力。
原文摘要 · Abstract (English)
The increasing global prevalence of mental disorders, such as depression and PTSD, requires objective and scalable diagnostic tools. Traditional clinical assessments often face limitations in accessibility, objectivity, and consistency. This paper investigates the potential of multimodal machine learning to address these challenges, leveraging the complementary information available in text, audio, and video data. Our approach involves a comprehensive analysis of various data preprocessing techniques, including novel chunking and utterance-based formatting strategies. We systematically evaluate a range of state-of-the-art embedding models for each modality and employ Convolutional Neural Networks (CNNs) and Bidirectional LSTM Networks (BiLSTMs) for feature extraction. We explore data-level, feature-level, and decision-level fusion techniques, including a novel integration of Large Language Model (LLM) predictions. We also investigate the impact of replacing Multilayer Perceptron classifiers with Support Vector Machines. We extend our analysis to severity prediction using PHQ-8 and PCL-C scores and multi-class classification (considering co-occurring conditions). Our results demonstrate that utterance-based chunking significantly improves performance, particularly for text and audio modalities. Decision-level fusion, incorporating LLM predictions, achieves the highest accuracy, with a balanced accuracy of 94.8% for depression and 96.2% for PTSD detection. The combination of CNN-BiLSTM architectures with utterance-level chunking, coupled with the integration of external LLM, provides a powerful and nuanced approach to the detection and assessment of mental health conditions. Our findings highlight the potential of MMML for developing more accurate, accessible, and personalized mental healthcare tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。