首个阿拉伯语心理健康帖子大规模自动标注数据集,助力早期筛查。
CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
- 用自动化方法构建阿拉伯语Reddit帖子数据集
- 涵盖6种心理疾病与对照组,规模和多样性领先现有资源
- 适合研究低资源语言心理健康检测的学者使用
全球数百万人都受心理健康问题困扰,但早期发现仍面临挑战,尤其在阿拉伯语群体中,受限于资源匮乏及文化对心理话题的回避。尽管英语研究丰富,阿拉伯语相关数据仍严重不足,主要因缺乏标注数据集。本文提出CARMA,首个基于自动化标注的阿拉伯语Reddit心理健康数据集,涵盖焦虑、自闭症、抑郁等六类心理状况及对照组。该数据集在规模与多样性上超越现有资源。我们通过定性与定量分析,揭示不同用户间词汇与语义差异,识别特定心理状态的语言特征。为验证数据集潜力,我们采用多种模型(从浅层分类器到大语言模型)进行分类实验,结果表明其在提升阿拉伯语等低资源语言的心理健康检测方面具有显著前景。
原文摘要 · Abstract (English)
Mental health disorders affect millions worldwide, yet early detection remains a major challenge, particularly for Arabic-speaking populations where resources are limited and mental health discourse is often discouraged due to cultural stigma. While substantial research has focused on English-language mental health detection, Arabic remains significantly underexplored, partly due to the scarcity of annotated datasets. We present CARMA, the first automatically annotated large-scale dataset of Arabic Reddit posts. The dataset encompasses six mental health conditions, such as Anxiety, Autism, and Depression, and a control group. CARMA surpasses existing resources in both scale and diversity. We conduct qualitative and quantitative analyses of lexical and semantic differences between users, providing insights into the linguistic markers of specific mental health conditions. To demonstrate the dataset's potential for further mental health analysis, we perform classification experiments using a range of models, from shallow classifiers to large language models. Our results highlight the promise of advancing mental health detection in underrepresented languages such as Arabic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。