小模型在急诊决策支持中表现优于大模型,无需医学微调即可高效工作。
Small Language Models for Emergency Departments Decision Support: A Benchmark Study
- 构建急诊场景专用评测集,测试小模型在医疗与通用任务上的表现。
- 通用小模型在多个医疗数据集上超越医学微调模型,准确率提升显著。
- 适合资源受限、注重隐私的医院实时决策系统部署。
大型语言模型(LLMs)在医疗领域日益流行,用于辅助医生完成各类临床与运营任务。急诊科环境快节奏且高风险,小语言模型(SLMs)因其参数量少、推理能力强、运行高效,具备显著潜力,可及时提供精准信息整合,提升临床决策与流程效率。本文提出一个综合性基准,旨在筛选适用于急诊决策支持的SLMs,兼顾专业医疗知识与广泛问题求解能力。评估聚焦于在通用与医学语料混合训练的小模型。强调使用SLMs的关键原因在于实际部署中的硬件限制、运营成本和隐私顾虑。评测数据集包括MedMCQA、MedQA-4Options和PubMedQA,医学摘要数据集模拟急诊医生日常任务。实验结果表明,通用领域训练的SLMs在多项基准上显著优于医学微调版本,说明在急诊场景中,模型无需专门医学微调即可达到良好效果。
原文摘要 · Abstract (English)
Large language models (LLMs) have become increasingly popular in medical domains to assist physicians with a variety of clinical and operational tasks. Given the fast-paced and high-stakes environment of emergency departments (EDs), small language models (SLMs), characterized by a reduction in parameter count compared to LLMs, offer significant potential due to their inherent reasoning capability and efficient performance. This enables SLMs to support physicians by providing timely and accurate information synthesis, thereby improving clinical decision-making and workflow efficiency. In this paper, we present a comprehensive benchmark designed to identify SLMs suited for ED decision support, taking into account both specialized medical expertise and broad general problem-solving capabilities. In our evaluations, we focus on SLMs that have been trained on a mixture of general-domain and medical corpora. A key motivation for emphasizing SLMs is the practical hardware limitations, operational cost constraints, and privacy concerns in the typical real-world deployments. Our benchmark datasets include MedMCQA, MedQA-4Options, and PubMedQA, with the medical abstracts dataset emulating tasks aligned with real ED physicians' daily tasks. Experimental results reveal that general-domain SLMs surprisingly outperform their medically fine-tuned counterparts across these diverse benchmarks for ED. This indicates that for ED, specialized medical fine-tuning of the model may not be required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。