arXiv:2504.04891cs.CLcs.LG2025-04被引 3

用大模型低成本实现多语种抑郁检测,深求V3表现最佳

Leveraging Large Language Models for Cost-Effective, Multilingual Depression Detection and Severity Assessment

  • 选用4个大模型对比,用临床访谈文本评估抑郁
  • 深求V3在零样本下准确率高,复杂场景AUC稳定
  • 适合临床筛查应用,但严重程度评估需改进

抑郁症是一种常见且难以早期发现的心理障碍,常因症状主观性而延误。本文评估了四种大语言模型在抑郁检测中的表现,使用临床访谈数据进行测试。选出表现最优的模型后,进一步验证其在严重程度评估和知识增强场景下的能力。在包含6种精神疾病共51,074条陈述的数据集上评估复杂诊断场景的鲁棒性。结果表明,DeepSeek V3在零样本和少样本场景中均表现优异,其中零样本最为高效;在复杂诊断场景中保持稳定的高AUC值,具有强检测能力。然而,与人工评估相比,其严重程度判断一致性较低,尤其对轻度抑郁患者。研究证实DeepSeek V3在真实临床环境中开展文本抑郁检测的巨大潜力,但也强调需进一步优化严重程度评估并减少潜在偏见以提升临床可靠性。

原文摘要 · Abstract (English)

Depression is a prevalent mental health disorder that is difficult to detect early due to subjective symptom assessments. Recent advancements in large language models have offered efficient and cost-effective approaches for this objective. In this study, we evaluated the performance of four LLMs in depression detection using clinical interview data. We selected the best performing model and further tested it in the severity evaluation scenario and knowledge enhanced scenario. The robustness was evaluated in complex diagnostic scenarios using a dataset comprising 51074 statements from six different mental disorders. We found that DeepSeek V3 is the most reliable and cost-effective model for depression detection, performing well in both zero-shot and few-shot scenarios, with zero-shot being the most efficient choice. The evaluation of severity showed low agreement with the human evaluator, particularly for mild depression. The model maintains stably high AUCs for detecting depression in complex diagnostic scenarios. These findings highlight DeepSeek V3s strong potential for text-based depression detection in real-world clinical applications. However, they also underscore the need for further refinement in severity assessment and the mitigation of potential biases to enhance clinical reliability.

抑郁检测大模型临床应用多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。