LLMs在临床预测上仍不如传统机器学习模型,需谨慎应用。
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
- 构建ClinicalBench基准,对比LLMs与传统模型在临床预测表现。
- 14个通用和8个医学LLM均未超越XGBoost等传统模型。
- 提醒医疗领域慎用LLMs,其临床推理能力仍有不足。
大型语言模型(LLMs)在医学文本处理和执业医师考试中展现出强大能力,但当前临床预测任务仍主要依赖支持向量机(SVM)和梯度提升树(XGBoost)等传统机器学习模型。本文提出一个新基准ClinicalBench,用于全面评估通用及医学专用LLMs在临床预测中的表现,并与11种传统模型对比。该基准涵盖3类常见临床预测任务、2个数据库、14个通用LLM、8个医学LLM和11种传统模型。通过大量实证研究发现,无论模型规模、提示策略或微调方式如何,通用和医学LLMs尚未在临床预测中超越传统模型,揭示其在临床推理与决策方面可能存在局限。研究呼吁在医疗应用中谨慎采用LLMs,同时强调ClinicalBench可作为连接模型发展与真实临床实践的桥梁。
原文摘要 · Abstract (English)
Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text processing tasks and medical licensing exams. Meanwhile, traditional ML models such as SVM and XGBoost have still been mainly adopted in clinical prediction tasks. An emerging question is: Can LLMs beat traditional ML models in clinical prediction? Thus, we build a new benchmark ClinicalBench to comprehensively study the clinical predictive modeling capacities of both general-purpose and medical LLMs, and compare them with traditional ML models. ClinicalBench embraces three common clinical prediction tasks, two databases, 14 general-purpose LLMs, 8 medical LLMs, and 11 traditional ML models. Through extensive empirical investigation, we discover that both general-purpose and medical LLMs, even with different model scales, diverse prompting or fine-tuning strategies, still cannot beat traditional ML models in clinical prediction yet, shedding light on their potential deficiency in clinical reasoning and decision-making. We call for caution when practitioners adopt LLMs in clinical applications. ClinicalBench can be utilized to bridge the gap between LLMs' development for healthcare and real-world clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。