用机器学习分析库尔德语推文,识别抑郁迹象。
Mining Mental Health Signals: A Comparative Study of Four Machine Learning Methods for Depression Detection from Social Media Posts in Sorani Kurdish
- 基于专家词表收集960条库尔德语推文,标注三类情绪状态。
- 随机森林模型准确率与F1值达80%,表现最佳。
- 首个针对索拉尼库尔德语抑郁检测的研究,适合语言处理初学者参考。
抑郁症是一种常见心理疾病,可能导致绝望、兴趣丧失、自残甚至自杀。由于患者常不主动报告或寻求治疗,早期发现困难。随着社交媒体普及,用户在线表达情绪日益频繁,为通过文本分析实现早期检测提供了新路径。尽管已有研究聚焦英语等语言,但针对索拉尼库尔德语的研究尚属空白。本文提出一种结合机器学习与自然语言处理的方法,用于检测索拉尼语推文中的抑郁信号。研究团队在专家指导下构建抑郁相关关键词集,从X(原Twitter)平台获取960条公开推文,并由大学医学院学者与高年级学生标注为三类:显示抑郁、未显示抑郁和可疑。训练并评估了四种监督模型——支持向量机、多项式朴素贝叶斯、逻辑回归与随机森林,其中随机森林在准确率与F1分数上均达到80%,表现最优。本研究为库尔德语语境下的自动化抑郁检测建立了基准。
原文摘要 · Abstract (English)
Depression is a common mental health condition that can lead to hopelessness, loss of interest, self-harm, and even suicide. Early detection is challenging due to individuals not self-reporting or seeking timely clinical help. With the rise of social media, users increasingly express emotions online, offering new opportunities for detection through text analysis. While prior research has focused on languages such as English, no studies exist for Sorani Kurdish. This work presents a machine learning and Natural Language Processing (NLP) approach to detect depression in Sorani tweets. A set of depression-related keywords was developed with expert input to collect 960 public tweets from X (Twitter platform). The dataset was annotated into three classes: Shows depression, Not-show depression, and Suspicious by academics and final year medical students at the University of Kurdistan Hewlêr. Four supervised models, including Support Vector Machines, Multinomial Naive Bayes, Logistic Regression, and Random Forest, were trained and evaluated, with Random Forest achieving the highest performance accuracy and F1-score of 80%. This study establishes a baseline for automated depression detection in Kurdish language contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。