首份非社交媒体心理疾病数据集综述,揭示研究空白与改进方向
Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

- 采用PRISMA方法系统梳理多语言非社交文本数据集
- 现有数据集多为英文,集中于抑郁检测且人群覆盖不均
- 适合心理健康研究者构建更可靠、多元的临床数据资源
及时发现心理疾病是重要的社会挑战。目前用于辅助检测的自然语言处理与机器学习方法主要依赖社交媒体数据,但此类数据常存在采样偏差及伦理隐私问题。一种突破路径是使用非社交媒体数据。本文首次全面回顾了适用于心理疾病研究的非社交媒体自由文本数据集,采用PRISMA方法进行系统调查,涵盖多种语言的数据集。研究发现,现有非社交文本数据集主要集中于英语,且以抑郁检测为主;同时在人口统计、平台来源、数据类型、标注方式和研究方法上差异显著。该综述揭示了关键缺口,并指出开发更具多样性、可靠性与临床相关性的数据资源的机遇。
原文摘要 · Abstract (English)
Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA methodology to conduct our survey and we review datasets available in multiple languages. We find that non-social media free-text based datasets are predominantly focused on English and on detecting depression. These datasets also vary in demographics, platforms, data types, annotation techniques, and methodologies. This systematic review also reveals key gaps and highlights opportunities to develop more diverse, reliable and clinically-relevant resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。