本地部署的小模型通过微调可实现急诊决策支持的临床竞争力。
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
- 用LoRA等策略微调开源小模型,实现本地化部署
- 在分诊和转诊任务上超越商用模型,诊断仍具挑战
- 能发现商用模型遗漏的重症患者,适合急诊场景
将大语言模型用于急诊科决策支持面临两大挑战:向封闭式商业模型传输患者数据存在隐私风险,且缺乏对本地部署开源小语言模型(SLMs)微调策略的系统评估。本研究在三个急诊任务上对比了八种开源SLMs,采用零样本提示、前缀微调、低秩适应(LoRA)和全量微调方法,使用2,083例MIMIC-IV-ED病例,并以Claude Haiku 4.5和Claude Sonnet 4.5为基准。结果表明,经LoRA微调的开源SLMs在分诊等级预测和专科转诊推荐任务上表现优于商用基线,而诊断预测仍具挑战;混淆矩阵分析显示,微调后的开源模型可识别出商用模型遗漏的高危患者。这些结果证明,本地部署的SLMs可在急诊决策支持中达到临床可用水平。
原文摘要 · Abstract (English)
Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs). We benchmarked eight open-source SLMs using zero-shot prompting, prefix tuning, Low-Rank Adaptation (LoRA), and full fine-tuning on three ED tasks: triage level prediction, specialist referral recommendation, and diagnosis prediction. Using 2,083 MIMIC-IV-ED cases and Claude Haiku 4.5 and Claude Sonnet 4.5 as baselines, we found that LoRA fine-tuned open-source SLMs outperform commercial baselines on triage level prediction and specialist referral recommendation, while diagnosis prediction remains challenging for open-source SLMs. Confusion matrix analysis further shows that fine-tuned open-source SLMs can detect highest-severity patients missed by the commercial baselines. These results demonstrate that locally deployable SLMs can achieve clinically competitive performance for ED decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。