arXiv:2509.26351cs.LG2025-09中稿 · the IEEE Internati…

用大模型构建可复现的急诊分诊基准,模拟真实医院与战场场景。

LLM-Assisted Emergency Triage Benchmark: Bridging Hospital-Rich and MCI-Like Field Simulation

  • 用大模型自动清洗和对齐急诊数据,降低使用门槛。
  • 在医院全数据与战场简化数据下均实现准确预测,但后者性能下降23%。
  • 提供可解释性分析,适合临床AI研究者和医疗系统开发者。

急诊与大规模伤亡事件(MCI)分诊研究受限于缺乏公开可复现的基准。这类场景需快速识别最危重患者,准确预测恶化可指导及时干预。尽管MIMIC-IV-ED数据库对认证研究者开放,但转化为分诊基准需大量预处理、特征对齐与模式整合,仅限技术背景强的用户使用。本文提出一个开源、基于大语言模型(LLM)的急诊分诊基准,用于预测重症监护室转入与院内死亡。该基准定义两种场景:(i) 医院丰富环境,包含生命体征、检验结果、病历记录、主诉和结构化观察;(ii) 类似战场的现场模拟,仅含生命体征、观察与病历。LLM直接参与数据构建,包括:(i) 清洗如AVPU、呼吸设备等噪声字段;(ii) 优先筛选临床相关生命体征与检验指标;(iii) 指导模式对齐与多表高效合并。我们还提供基线模型与基于SHAP的可解释性分析,揭示两类场景间的预测差距及关键特征。这些工作使分诊预测研究更可复现、更易获取,推动临床AI数据民主化。

原文摘要 · Abstract (English)

Research on emergency and mass casualty incident (MCI) triage has been limited by the absence of openly usable, reproducible benchmarks. Yet these scenarios demand rapid identification of the patients most in need, where accurate deterioration prediction can guide timely interventions. While the MIMIC-IV-ED database is openly available to credentialed researchers, transforming it into a triage-focused benchmark requires extensive preprocessing, feature harmonization, and schema alignment -- barriers that restrict accessibility to only highly technical users. We address these gaps by first introducing an open, LLM-assisted emergency triage benchmark for deterioration prediction (ICU transfer, in-hospital mortality). The benchmark then defines two regimes: (i) a hospital-rich setting with vitals, labs, notes, chief complaints, and structured observations, and (ii) an MCI-like field simulation limited to vitals, observations, and notes. Large language models (LLMs) contributed directly to dataset construction by (i) harmonizing noisy fields such as AVPU and breathing devices, (ii) prioritizing clinically relevant vitals and labs, and (iii) guiding schema alignment and efficient merging of disparate tables. We further provide baseline models and SHAP-based interpretability analyses, illustrating predictive gaps between regimes and the features most critical for triage. Together, these contributions make triage prediction research more reproducible and accessible -- a step toward dataset democratization in clinical AI.

急诊分诊大模型临床AI数据基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。