用大模型提升精神科急诊再入院预测的准确性和医生可理解性。
Explainable AI for Mental Health Emergency Returns: Integrating LLMs with Predictive Modeling
- 用大模型自动提取病历关键信息,增强传统模型输入。
- 预测准确率提升至AUC 0.76,解释生成准确率达99%。
- 适合临床医生需要快速理解模型决策的场景。
精神科急诊再入院是重大医疗负担,30天内再入院率高达24%-27%。传统机器学习模型缺乏临床可解释性。本研究回顾分析了2018年1月至2022年12月间深南部某医学院附属医院42,464次精神科急诊就诊记录(涉及27,904名患者)。主要评估两个指标:30天再入院预测准确率与基于新型LLM增强框架的模型可解释性,该框架结合SHAP值与临床知识。在主诉分类中,使用10样本学习的LLaMA 3(8B)模型表现最优,准确率为0.882,F1得分为0.86;在社会决定因素(SDoH)分类中,大模型实现0.95准确率和0.96 F1得分,其中酒精、烟草及物质滥用类别表现最佳(F1: 0.96-0.89),而运动与居住环境类别较低(F1: 0.70-0.67)。LLM增强的解释框架将预测结果转化为临床相关解释的准确率达99%。利用大模型提取特征后,XGBoost模型的AUC从0.74提升至0.76,AUC-PR从0.58升至0.61。结论表明,融合大模型与机器学习能带来稳定但适度的性能提升,并显著改善模型可解释性,为将预测分析转化为临床可操作洞察提供可行框架。
原文摘要 · Abstract (English)
Importance: Emergency department (ED) returns for mental health conditions pose a major healthcare burden, with 24-27% of patients returning within 30 days. Traditional machine learning models for predicting these returns often lack interpretability for clinical use. Objective: To assess whether integrating large language models (LLMs) with machine learning improves predictive accuracy and clinical interpretability of ED mental health return risk models. Methods: This retrospective cohort study analyzed 42,464 ED visits for 27,904 unique mental health patients at an academic medical center in the Deep South from January 2018 to December 2022. Main Outcomes and Measures: Two primary outcomes were evaluated: (1) 30-day ED return prediction accuracy and (2) model interpretability using a novel LLM-enhanced framework integrating SHAP (SHapley Additive exPlanations) values with clinical knowledge. Results: For chief complaint classification, LLaMA 3 (8B) with 10-shot learning outperformed traditional models (accuracy: 0.882, F1-score: 0.86). In SDoH classification, LLM-based models achieved 0.95 accuracy and 0.96 F1-score, with Alcohol, Tobacco, and Substance Abuse performing best (F1: 0.96-0.89), while Exercise and Home Environment showed lower performance (F1: 0.70-0.67). The LLM-based interpretability framework achieved 99% accuracy in translating model predictions into clinically relevant explanations. LLM-extracted features improved XGBoost AUC from 0.74 to 0.76 and AUC-PR from 0.58 to 0.61. Conclusions and Relevance: Integrating LLMs with machine learning models yielded modest but consistent accuracy gains while significantly enhancing interpretability through automated, clinically relevant explanations. This approach provides a framework for translating predictive analytics into actionable clinical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。