arXiv:2410.03908cs.CLcs.AI2024-10EMNLP被引 17

首个针对抑郁焦虑共病的社交平台诊断基准,揭示大模型在复杂心理诊断中的局限。

Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis

  • 构建多标签标注数据集ANGST,支持抑郁与焦虑共现识别
  • 大模型最高F1仅72%,显示当前技术仍难胜任复杂共病诊断
  • 适合心理健康AI研究者及临床辅助系统开发者参考

本研究提出ANGST,首个针对社交媒体中抑郁-焦虑共病分类的基准数据集。不同于以往将精神疾病视为孤立状况的简化处理,ANGST支持多标签分类,允许每条帖子同时标记为抑郁和/或焦虑。数据集包含2876条由专业心理学家精标的内容,以及额外7667条银标数据,更真实反映线上心理讨论。我们使用Mental-BERT至GPT-4等前沿语言模型对ANGST进行评测。结果表明,尽管GPT-4表现最优,但在多类别共病分类任务中,所有模型的F1分数均未超过72%,凸显当前语言模型在复杂心理诊断中的显著挑战。

原文摘要 · Abstract (English)

In this study, we introduce ANGST, a novel, first-of-its kind benchmark for depression-anxiety comorbidity classification from social media posts. Unlike contemporary datasets that often oversimplify the intricate interplay between different mental health disorders by treating them as isolated conditions, ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety. Comprising 2876 meticulously annotated posts by expert psychologists and an additional 7667 silver-labeled posts, ANGST posits a more representative sample of online mental health discourse. Moreover, we benchmark ANGST using various state-of-the-art language models, ranging from Mental-BERT to GPT-4. Our results provide significant insights into the capabilities and limitations of these models in complex diagnostic scenarios. While GPT-4 generally outperforms other models, none achieve an F1 score exceeding 72% in multi-class comorbid classification, underscoring the ongoing challenges in applying language models to mental health diagnostics.

心理AI共病诊断大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。