构建首个统一的社交媒体心理健康检测数据集套件,支持多任务研究。
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection

- 基于Reddit构建四类互补任务数据集,标注严格且一致
- 模型在各任务上F1达93%-99%,验证数据高质量
- 适合心理健康NLP研究者做基准测试与模型对比
在线支持社区的兴起为通过自然语言处理(NLP)研究心理健康提供了新视角。然而,高质量、经过充分验证的数据集匮乏,制约了该领域发展。现有研究多构建特定任务语料库,缺乏共享资源,导致复现困难与跨任务比较难以实现。本文提出一个统一的基准数据集套件,包含四个基于Reddit的互补任务:(i) 自杀意念检测,(ii) 二分类一般心理障碍检测,(iii) 双相情感障碍检测,(iv) 多分类心理障碍识别。所有数据集均基于严谨的语言学审查、明确的标注指南和人工验证建立,标注者间一致性指标始终超过0.8的基线水平,确保标签可信。先前工作在Transformer和上下文感知循环模型上的表现显示,这些模型在各项任务中取得了93%-99%的F1分数,进一步验证了数据集的有效性。该套件为可复现的心理健康NLP研究提供统一基础,支持跨任务基准测试、多任务学习与公平模型比较。本研究向学术界提供了一个易于获取、多样化的心智健康计算研究资源。
原文摘要 · Abstract (English)
The growing availability of online support groups has opened up new windows to study mental health through natural language processing (NLP). However, it is hindered by a lack of high-quality, well-validated datasets. Existing studies have a tendency to build task-specific corpora without collecting them into widely available resources, and this makes reproducibility as well as cross-task comparison challenging. In this paper, we present a uniform benchmark set of four Reddit-based datasets for disjoint but complementary tasks: (i) detection of suicidal ideation, (ii) binary general mental disorder detection, (iii) bipolar disorder detection, and (iv) multi-class mental disorder classification. All datasets were established upon diligent linguistic inspection, well-defined annotation guidelines, and human-judgmental verification. Inter-annotator agreement metrics always exceeded the baseline agreement score of 0.8, ensuring the labels' trustworthiness. Previous work's evidence of performance on both transformer and contextualized recurrent models demonstrates that these models receive excellent performances on tasks (F1 ~ 93-99%), further validating the usefulness of the datasets. By combining these resources, we establish a unifying foundation for reproducible mental health NLP studies with the ability to carry out cross-task benchmarking, multi-task learning, and fair model comparison. The presented benchmark suite provides the research community with an easy-to-access and varied resource for advancing computational approaches toward mental health research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。