arXiv:2409.13713cs.CLcs.LG2024-09被引 24

用情感特征提升抑郁症早期检测准确率

Sentiment Informed Sentence BERT-Ensemble Algorithm for Depression Detection

  • 将情感指标作为额外特征融入SBERT集成模型
  • 在两个数据集上分别达到69%和76%的F1分数
  • 适合心理健康与NLP交叉研究者参考

世界卫生组织(WHO)指出全球约有2.8亿人患有抑郁症。然而,利用机器学习技术进行早期抑郁症检测的研究仍有限。以往研究多采用单一算法,难以应对数据复杂性,易过拟合且泛化能力差。本文基于两个基准社交媒体数据集(D1 和 D2),评估了多种机器学习算法在早期抑郁症检测中的表现,并引入情感指标以提升模型性能。实验结果表明,将句子双向编码器表示(SBERT)的数值向量输入堆叠集成模型,在数据集D1上获得69%的F1分数,在数据集D2上达到76%。研究结果表明,将情感指标作为附加特征可有效提升检测性能,建议未来构建抑郁相关词汇语料库以进一步优化模型。

原文摘要 · Abstract (English)

The World Health Organisation (WHO) revealed approximately 280 million people in the world suffer from depression. Yet, existing studies on early-stage depression detection using machine learning (ML) techniques are limited. Prior studies have applied a single stand-alone algorithm, which is unable to deal with data complexities, prone to overfitting, and limited in generalization. To this end, our paper examined the performance of several ML algorithms for early-stage depression detection using two benchmark social media datasets (D1 and D2). More specifically, we incorporated sentiment indicators to improve our model performance. Our experimental results showed that sentence bidirectional encoder representations from transformers (SBERT) numerical vectors fitted into the stacking ensemble model achieved comparable F1 scores of 69% in the dataset (D1) and 76% in the dataset (D2). Our findings suggest that utilizing sentiment indicators as an additional feature for depression detection yields an improved model performance, and thus, we recommend the development of a depressive term corpus for future work.

抑郁症检测情感分析SBERT集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。