arXiv:2604.19887cs.CLcs.AI2026-04

用大模型分析社交媒体文本,自动评估抑郁风险。

Depression Risk Assessment in Social Media via Large Language Models

论文配图:Depression Risk Assessment in Social Media via Large Language Models
图 1 · 摘自论文原文
  • 通过多标签分类识别8种抑郁相关情绪,计算加权严重度指数。
  • 零样本下在6000条数据上达到微平均F1 0.75,接近精调模型表现。
  • 适用于大规模心理状态监测,尤其适合研究社群差异的学者。

抑郁症是全球最普遍且最具破坏性的心理健康问题之一,常被漏诊和忽视。社交媒体平台提供了丰富的自然语言信号,可用于自动化监测心理状态。本文提出一种基于大语言模型(LLM)的系统,针对Reddit帖子进行抑郁风险评估,通过多标签分类识别八种与抑郁相关的负面情绪,并计算加权严重度指数。该方法在标注的DepressionEmo数据集(约6,000条帖子)上以零样本方式评估,同时应用于2024–2025年间从四个子论坛收集的469,692条评论。最佳模型gemma3:27b实现微平均F1为0.75,宏平均F1为0.70,性能可与专门微调的模型(BART:微平均F1 0.80,宏平均F1 0.76)媲美。在真实世界应用中,不同社区表现出稳定且一致的风险特征,其中r/depression与r/anxiety群体间差异显著。研究证明了低成本、可扩展的大规模心理监测方法可行性。

原文摘要 · Abstract (English)

Depression is one of the most prevalent and debilitating mental health conditions worldwide, frequently underdiagnosed and undertreated. The proliferation of social media platforms provides a rich source of naturalistic linguistic signals for the automated monitoring of psychological well-being. In this work, we propose a system based on Large Language Models (LLMs) for depression risk assessment in Reddit posts, through multi-label classification of eight depression-associated emotions and the computation of a weighted severity index. The method is evaluated in a zero-shot setting on the annotated DepressionEmo dataset (~6,000 posts) and applied in-the-wild to 469,692 comments collected from four subreddits over the period 2024-2025. Our best model, gemma3:27b, achieves micro-F1 = 0.75 and macro-F1 = 0.70, results competitive with purpose-built fine-tuned models (BART: micro-F1 = 0.80, macro-F1 = 0.76). The in-the-wild analysis reveals consistent and temporally stable risk profiles across communities, with marked differences between r/depression and r/anxiety. Our findings demonstrate the feasibility of a cost-effective, scalable approach for large-scale psychological monitoring.

抑郁评估大模型社交媒体情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。