医学领域微调对大模型表现提升有限,多数反不如基础模型。
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
- 直接对比医疗模型与基础模型,避免误导性评估
- 3次提示下仅26.7%场景优于基线,56.7%反而更差
- 强调统计显著性,适合关注模型真实能力的研究者
近期多项研究通过在公开生物医学语料上继续预训练,尝试将通用大语言模型(LLMs)和视觉-语言模型(VLMs)适配至医疗场景。本文对比了十种医疗专用LLM和两种医疗专用VLM与其对应基础模型的表现,发现几乎所有医疗模型在零样本/少样本提示和监督微调条件下,未能持续优于基础模型。例如,在基于临床笔记的问答任务中,3次提示设置下,医疗LLM仅在26.7%情况下表现更优,16.7%持平,56.7%显著更差。结论基于三点:(i)每种医疗模型与其基线直接对比;(ii)为每模型独立优化提示词;(iii)考虑统计不确定性。结果表明,当前顶尖通用模型可能已具备较强的医学知识与推理能力,对后续研究提出严谨评估建议。
原文摘要 · Abstract (English)
Several recent works seek to adapt general-purpose large language models (LLMs) and vision-language models (VLMs) for medical applications through continued pretraining on publicly available biomedical corpora. These works typically claim that such domain-adaptive pretraining improves performance on various downstream medical tasks, such as answering medical exam questions. In this paper, we compare ten "medical" LLMs and two VLMs against their corresponding base models, arriving at a different conclusion: all medical VLMs and nearly all medical LLMs fail to consistently improve over their base models in the zero-/few-shot prompting and supervised fine-tuning regimes for medical question answering (QA). For instance, on clinical-note-based QA tasks in the 3-shot setting, medical LLMs outperform their base models in only 26.7% of cases, reach a (statistical) tie in 16.7% of cases, and perform significantly worse in the remaining 56.7% of cases. Our conclusions are based on (i) comparing each medical model directly against its base model; (ii) optimizing the prompts for each model separately in zero-/few-shot prompting; and (iii) accounting for statistical uncertainty in comparisons. Our findings suggest that state-of-the-art general-domain models may already exhibit strong medical knowledge and reasoning capabilities, and offer recommendations to strengthen the conclusions of future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。