专为医院运营训练的模型比通用模型更准,关键在真实病历数据预训练和微调。
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
- 用800亿临床文本+6270亿网络文本预训练,专注医院运营任务。
- 微调后模型性能超大通用模型70倍,读卡率提升23.66%。
- 适合医疗AI研发者、医院系统优化者,需真实病历数据支持。
医院运营依赖患者流转、成本与照护质量等决策。尽管通用基础模型在医学知识和对话任务上表现良好,但在临床运营决策方面仍缺乏专门知识。我们提出Lang1系列模型(100M-7B参数),在800亿条来自纽约大学朗格尼健康中心电子病历(EHR)的临床文本与6270亿条互联网文本上进行预训练。为真实评估其表现,我们构建了基于668,331份病历的REalistic Medical Evaluation(ReMedE)基准,涵盖五个关键任务:30天再入院预测、30天死亡率预测、住院时长、共病编码及保险拒赔预测。零样本测试中,通用与专用模型在四项任务上表现不佳(AUROC 36.6%-71.7%),仅死亡率预测例外。微调后,Lang1-1B在性能上超越70倍更大的通用模型和671倍更大的零样本模型,分别提升AUROC 3.64%-6.75%与1.66%-23.66%。联合多任务微调还带来跨任务提升。Lang1-1B在分布外场景(如其他临床任务、外部医疗系统)中也表现稳健。研究表明,医院运营预测需显式监督微调,而领域内预训练可显著提升效率。这支持了专业化大模型在特定任务上可媲美通用模型的观点,并强调有效医疗系统AI需结合领域预训练、监督微调与真实世界评估。
原文摘要 · Abstract (English)
Hospitals and healthcare systems rely on operational decisions that determine patient flow, cost, and quality of care. Despite strong performance on medical knowledge and conversational benchmarks, foundation models trained on general text may lack the specialized knowledge required for these operational decisions. We introduce Lang1, a family of models (100M-7B parameters) pretrained on a specialized corpus blending 80B clinical tokens from NYU Langone Health's EHRs and 627B tokens from the internet. To rigorously evaluate Lang1 in real-world settings, we developed the REalistic Medical Evaluation (ReMedE), a benchmark derived from 668,331 EHR notes that evaluates five critical tasks: 30-day readmission prediction, 30-day mortality prediction, length of stay, comorbidity coding, and predicting insurance claims denial. In zero-shot settings, both general-purpose and specialized models underperform on four of five tasks (36.6%-71.7% AUROC), with mortality prediction being an exception. After finetuning, Lang1-1B outperforms finetuned generalist models up to 70x larger and zero-shot models up to 671x larger, improving AUROC by 3.64%-6.75% and 1.66%-23.66% respectively. We also observed cross-task scaling with joint finetuning on multiple tasks leading to improvement on other tasks. Lang1-1B effectively transfers to out-of-distribution settings, including other clinical tasks and an external health system. Our findings suggest that predictive capabilities for hospital operations require explicit supervised finetuning, and that this finetuning process is made more efficient by in-domain pretraining on EHR. Our findings support the emerging view that specialized LLMs can compete with generalist models in specialized tasks, and show that effective healthcare systems AI requires the combination of in-domain pretraining, supervised finetuning, and real-world evaluation beyond proxy benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。