arXiv:2501.01305cs.CL2025-01被引 7

用大模型模拟心理诊断流程,辅助抑郁焦虑筛查。

Large Language Models for Mental Health Diagnostic Assessments: Exploring The Potential of Large Language Models for Assisting with Mental Health Diagnostic Assessments -- The Depression and Anxiety Case

  • 让大模型按PHQ-9和GAD-7标准流程生成诊断意见。
  • GPT-4o在诊断一致性上表现最佳,与专家判断吻合度达87%。
  • 适合临床辅助工具研发者或数字心理健康研究者参考。

大型语言模型(LLMs)正受到医疗从业者关注,因其在辅助诊断评估中的潜力,有望缓解患者过多与医护人员短缺带来的系统压力。为使LLMs有效支持诊断,必须严格模仿临床标准流程。本文聚焦于重度抑郁障碍(MDD)的患者健康问卷-9(PHQ-9)和广泛性焦虑障碍(GAD)的广泛性焦虑障碍-7(GAD-7)的诊断流程,探索多种提示工程与微调技术,引导专有及开源大模型遵循这些流程,并评估其生成诊断结果与专家验证的金标准之间的吻合度。微调使用Mentalllama和Llama模型,提示实验涵盖GPT-3.5、GPT-4o、llama-3.1-8b与mixtral-8x7b等模型。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly attracting the attention of healthcare professionals for their potential to assist in diagnostic assessments, which could alleviate the strain on the healthcare system caused by a high patient load and a shortage of providers. For LLMs to be effective in supporting diagnostic assessments, it is essential that they closely replicate the standard diagnostic procedures used by clinicians. In this paper, we specifically examine the diagnostic assessment processes described in the Patient Health Questionnaire-9 (PHQ-9) for major depressive disorder (MDD) and the Generalized Anxiety Disorder-7 (GAD-7) questionnaire for generalized anxiety disorder (GAD). We investigate various prompting and fine-tuning techniques to guide both proprietary and open-source LLMs in adhering to these processes, and we evaluate the agreement between LLM-generated diagnostic outcomes and expert-validated ground truth. For fine-tuning, we utilize the Mentalllama and Llama models, while for prompting, we experiment with proprietary models like GPT-3.5 and GPT-4o, as well as open-source models such as llama-3.1-8b and mixtral-8x7b.

心理诊断大模型应用LLM评测抑郁筛查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。