arXiv:2510.11233cs.CL2025-10被引 1

构建中文社交媒体抑郁风险检测数据集,支持细粒度心理分析

CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis

  • 收集233名用户4.4万条社交帖子,专家标注1万多段抑郁相关文本
  • 提供二分类标签与多维心理属性,支持抑郁信号的可解释分析
  • 适用于中文NLP任务,助力本土化心理健康应用开发

抑郁症是全球重要的公共健康问题,但现有的中文抑郁风险检测资源仍匮乏,且多集中于二分类任务。为弥补这一不足,我们发布CNSocialDepress——一个面向中文社交媒体的抑郁风险检测基准数据集。该数据集包含233名用户的44,178条帖子,心理专家标注了10,306段抑郁相关文本。CNSocialDepress不仅提供二分类风险标签,还包含结构化的多维度心理属性,支持对抑郁信号的可解释、细粒度分析。实验表明,该数据集在多种自然语言处理任务中具有实用价值,包括结构化心理画像构建及大语言模型微调以实现抑郁检测。全面评估显示其在抑郁风险识别与心理分析方面的有效性与实际应用潜力,为面向中文人群的心理健康应用提供重要支持。

原文摘要 · Abstract (English)

Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce and largely focus on binary classification. To address this limitation, we release CNSocialDepress, a benchmark dataset for depression risk detection on Chinese social media. The dataset contains 44,178 posts from 233 users; psychological experts annotated 10,306 depression-related segments. CNSocialDepress provides binary risk labels along with structured, multidimensional psychological attributes, enabling interpretable and fine-grained analyses of depressive signals. Experimental results demonstrate the dataset's utility across a range of NLP tasks, including structured psychological profiling and fine-tuning large language models for depression detection. Comprehensive evaluations highlight the dataset's effectiveness and practical value for depression risk identification and psychological analysis, thereby providing insights for mental health applications tailored to Chinese-speaking populations.

抑郁症检测中文数据集心理分析社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。