用大模型从患者消息中识别共病抑郁焦虑,效果超90%。
Optimizing Large Language Models for Detecting Symptoms of Comorbid Depression or Anxiety in Chronic Diseases: Insights from Patient Messages
- 通过提示工程与零样本学习优化模型,提升症状检测能力。
- 三款模型F1和准确率超90%,最强达93%。
- 适合用于慢病患者心理筛查系统,助力临床早干预。
糖尿病患者易伴发抑郁或焦虑,增加管理难度。本研究评估了大语言模型(LLMs)从安全患者消息中检测这些症状的表现。采用多种方法,包括提示工程、系统性角色设定、温度调节,以及零样本和少量样本学习,以确定最优模型并提升性能。五款模型中有三款表现优异(F-1与准确率均超90%),其中Llama 3.1 405B在零样本方法下达到93%的F-1和准确率。尽管模型在二分类任务及患者健康问卷-4等复杂指标上表现良好,但在复杂病例中仍存在不一致,需进一步开展真实场景评估。研究结果表明,LLMs有望辅助实现及时筛查与转诊,为实际分诊系统提供重要实证依据,有助于改善慢性病患者的心理健康服务。
原文摘要 · Abstract (English)
Patients with diabetes are at increased risk of comorbid depression or anxiety, complicating their management. This study evaluated the performance of large language models (LLMs) in detecting these symptoms from secure patient messages. We applied multiple approaches, including engineered prompts, systemic persona, temperature adjustments, and zero-shot and few-shot learning, to identify the best-performing model and enhance performance. Three out of five LLMs demonstrated excellent performance (over 90% of F-1 and accuracy), with Llama 3.1 405B achieving 93% in both F-1 and accuracy using a zero-shot approach. While LLMs showed promise in binary classification and handling complex metrics like Patient Health Questionnaire-4, inconsistencies in challenging cases warrant further real-life assessment. The findings highlight the potential of LLMs to assist in timely screening and referrals, providing valuable empirical knowledge for real-world triage systems that could improve mental health care for patients with chronic diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。