用大模型自动提取儿童抑郁症状,提升筛查效率与准确性。
LLM Assistance for Pediatric Depression
- 利用大模型零样本提取6-24岁儿童抑郁症状,替代传统问卷。
- Flan模型在罕见症状上表现优异,F1达0.92,整体精确率达0.78。
- 结果可作为机器学习特征,显著提升抑郁症识别准确率。
传统抑郁筛查工具如PHQ-9在儿科初级医疗中对儿童实施困难,受实际限制影响。人工智能有潜力缓解此问题,但精神健康领域标注数据稀缺且训练成本高,亟需高效零样本方法。本文研究了先进大模型在儿童(6-24岁)抑郁症状提取中的可行性,旨在辅助传统筛查并减少误诊。结果显示,所有大模型比关键词匹配方法效率高出60%,其中Flan在精度上领先(平均F1: 0.65,精确率: 0.78),在“睡眠问题”(F1: 0.92)和“自我厌恶”(F1: 0.8)等罕见症状上表现突出;Phi在精确率(0.44)与召回率(0.60)间取得平衡,在“情绪低落”(0.69)和“体重变化”(0.78)类别表现良好;Llama 3虽召回率最高(0.90),但存在过度泛化问题,不适用于此类分析。主要挑战包括临床笔记内容跨时间轨迹复杂及对PHQ-9得分的误读。最终验证,由Flan提供的症状标注作为机器学习特征,使抑郁症与对照组区分准确率达到0.78,远超无此特征的基线。
原文摘要 · Abstract (English)
Traditional depression screening methods, such as the PHQ-9, are particularly challenging for children in pediatric primary care due to practical limitations. AI has the potential to help, but the scarcity of annotated datasets in mental health, combined with the computational costs of training, highlights the need for efficient, zero-shot approaches. In this work, we investigate the feasibility of state-of-the-art LLMs for depressive symptom extraction in pediatric settings (ages 6-24). This approach aims to complement traditional screening and minimize diagnostic errors. Our findings show that all LLMs are 60% more efficient than word match, with Flan leading in precision (average F1: 0.65, precision: 0.78), excelling in the extraction of more rare symptoms like "sleep problems" (F1: 0.92) and "self-loathing" (F1: 0.8). Phi strikes a balance between precision (0.44) and recall (0.60), performing well in categories like "Feeling depressed" (0.69) and "Weight change" (0.78). Llama 3, with the highest recall (0.90), overgeneralizes symptoms, making it less suitable for this type of analysis. Challenges include the complexity of clinical notes and overgeneralization from PHQ-9 scores. The main challenges faced by LLMs include navigating the complex structure of clinical notes with content from different times in the patient trajectory, as well as misinterpreting elevated PHQ-9 scores. We finally demonstrate the utility of symptom annotations provided by Flan as features in an ML algorithm, which differentiates depression cases from controls with high precision of 0.78, showing a major performance boost compared to a baseline that does not use these features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。