arXiv:2507.19899cs.CL2025-07被引 3

构建首个抑郁症状细粒度标注数据集,评估大模型解释能力

A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs

  • 1017条社交媒体文本标注抑郁症状片段与12类症状
  • 发现GPT-4.1等模型在零样本/少样本下解释质量差异显著
  • 适合心理计算、可解释AI及临床辅助系统研究者

从社交媒体内容中早期检测抑郁可为心理健康干预提供及时支持。本文构建了一个高质量、专家标注的数据集,包含1,017条社交帖子,每条标注了抑郁症状片段,并映射到12种抑郁症状类别。不同于以往仅提供粗粒度帖子级标签的数据集,本数据集支持对模型预测和生成解释的细粒度评估。我们设计了一套评估框架,利用该临床基础数据集衡量大语言模型(LLMs)生成自然语言解释的忠实度与质量。通过精心设计的提示策略,包括零样本与少量样本方法及领域适配示例,评估了GPT-4.1、Gemini 2.5 Pro和Claude 3.7 Sonnet等主流专有模型。综合实证分析揭示这些模型在临床解释任务上的表现存在显著差异,且提示策略影响明显。研究强调了人类专业经验对引导大模型行为的重要性,为实现更安全、透明的心理健康人工智能系统迈出关键一步。

原文摘要 · Abstract (English)

Early detection of depression from online social media posts holds promise for providing timely mental health interventions. In this work, we present a high-quality, expert-annotated dataset of 1,017 social media posts labeled with depressive spans and mapped to 12 depression symptom categories. Unlike prior datasets that primarily offer coarse post-level labels \cite{cohan-etal-2018-smhd}, our dataset enables fine-grained evaluation of both model predictions and generated explanations. We develop an evaluation framework that leverages this clinically grounded dataset to assess the faithfulness and quality of natural language explanations generated by large language models (LLMs). Through carefully designed prompting strategies, including zero-shot and few-shot approaches with domain-adapted examples, we evaluate state-of-the-art proprietary LLMs including GPT-4.1, Gemini 2.5 Pro, and Claude 3.7 Sonnet. Our comprehensive empirical analysis reveals significant differences in how these models perform on clinical explanation tasks, with zero-shot and few-shot prompting. Our findings underscore the value of human expertise in guiding LLM behavior and offer a step toward safer, more transparent AI systems for psychological well-being.

抑郁症检测大模型解释数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。