arXiv:2507.01080cs.LGcs.PF2025-07被引 4

用大模型提升急诊分诊准确率,效果优于传统AI和人工。

Development and Comparative Evaluation of Three Artificial Intelligence Models (NLP, LLM, JEPA) for Predicting Triage in Emergency Departments: A 7-Month Retrospective Proof-of-Concept

  • 采用大语言模型处理原始病历与结构化数据,提取关键信息。
  • 大模型在预测住院需求上达到F1=0.900,AUC=0.879,最优。
  • 适合急诊科想提升分诊效率与安全性的医疗机构参考。

急诊科长期面临分诊错误问题,尤其在患者量大增与人力短缺背景下,漏诊和误诊更严重。本研究基于法国里尔罗杰·萨兰格罗医院7个月的成人分诊数据,评估了三种AI模型(TRIAGEMASTER,NLP;URGENTIAPARSE,LLM;EMERGINET,JEPA)对FRENCH分诊标准及护士实践的预测能力。其中,基于大语言模型的URGENTIAPARSE在所有指标中表现最佳,达最高准确率(F1-score 0.900,AUC-ROC 0.879),在预测住院需求(GEMSA)方面也显著领先。其在结构化数据与原始语音转录文本上的稳定表现,凸显了大语言模型在抽象患者信息方面的优势。结果表明,将大模型引入急诊流程可显著提升患者安全与运营效率,但实际应用仍需克服局限并保障伦理透明。

原文摘要 · Abstract (English)

Emergency departments struggle with persistent triage errors, especially undertriage and overtriage, which are aggravated by growing patient volumes and staff shortages. This study evaluated three AI models [TRIAGEMASTER (NLP), URGENTIAPARSE (LLM), and EMERGINET (JEPA)] against the FRENCH triage scale and nurse practice, using seven months of adult triage data from Roger Salengro Hospital in Lille, France. Among the models, the LLM-based URGENTIAPARSE consistently outperformed both AI alternatives and nurse triage, achieving the highest accuracy (F1-score 0.900, AUC-ROC 0.879) and superior performance in predicting hospitalization needs (GEMSA). Its robustness across structured data and raw transcripts highlighted the advantage of LLM architectures in abstracting patient information. Overall, the findings suggest that integrating LLM-based AI into emergency department workflows could significantly enhance patient safety and operational efficiency, though successful adoption will depend on addressing limitations and ensuring ethical transparency.

急诊分诊大语言模型AI医疗临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。