首个针对抑郁焦虑共病的社交平台诊断基准,揭示大模型在复杂心理诊断中的局限。
Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis
- 构建多标签标注数据集ANGST,支持抑郁与焦虑共现识别
- 大模型最高F1仅72%,显示当前技术仍难胜任复杂共病诊断
- 适合心理健康AI研究者及临床辅助系统开发者参考
本研究提出ANGST,首个针对社交媒体中抑郁-焦虑共病分类的基准数据集。不同于以往将精神疾病视为孤立状况的简化处理,ANGST支持多标签分类,允许每条帖子同时标记为抑郁和/或焦虑。数据集包含2876条由专业心理学家精标的内容,以及额外7667条银标数据,更真实反映线上心理讨论。我们使用Mental-BERT至GPT-4等前沿语言模型对ANGST进行评测。结果表明,尽管GPT-4表现最优,但在多类别共病分类任务中,所有模型的F1分数均未超过72%,凸显当前语言模型在复杂心理诊断中的显著挑战。
原文摘要 · Abstract (English)
In this study, we introduce ANGST, a novel, first-of-its kind benchmark for depression-anxiety comorbidity classification from social media posts. Unlike contemporary datasets that often oversimplify the intricate interplay between different mental health disorders by treating them as isolated conditions, ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety. Comprising 2876 meticulously annotated posts by expert psychologists and an additional 7667 silver-labeled posts, ANGST posits a more representative sample of online mental health discourse. Moreover, we benchmark ANGST using various state-of-the-art language models, ranging from Mental-BERT to GPT-4. Our results provide significant insights into the capabilities and limitations of these models in complex diagnostic scenarios. While GPT-4 generally outperforms other models, none achieve an F1 score exceeding 72% in multi-class comorbid classification, underscoring the ongoing challenges in applying language models to mental health diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。