arXiv:2411.04118cs.CLcs.AI2024-11EMNLP被引 50

医学大模型未必比通用模型强,实测多数表现更差。

Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

  • 直接对比医疗模型与基线模型,避免误导性评估
  • 3次提示下仅12.1%的医疗模型胜过基础模型
  • 强调严谨评估方法对结论的关键影响

近期多项研究通过在公开生物医学语料上继续预训练,开发面向医疗领域的大型语言模型(LLMs)和视觉-语言模型(VLMs)。这些工作普遍声称领域自适应预训练(DAPT)能提升下游医疗任务表现,如医学执照考试问答。本文将七个公开的“医疗”LLMs和两个VLMs与对应基线模型进行直接比较,发现:在零样本/少样本提示设置下,所有医疗VLMs及几乎全部医疗LLMs并未持续优于其基线模型。例如,在3次提示设置下,医疗LLMs仅在12.1%的情况下优于基线,49.8%持平,38.2%显著更差。结论基于三点:(i) 每个医疗模型与对应基线直接对比;(ii) 为每模型单独优化提示;(iii) 考虑统计不确定性。我们发现,当前通用领域模型可能已具备强大的医学知识与推理能力,建议未来研究采用更严谨评估方式。

原文摘要 · Abstract (English)

Several recent works seek to develop foundation models specifically for medical applications, adapting general-purpose large language models (LLMs) and vision-language models (VLMs) via continued pretraining on publicly available biomedical corpora. These works typically claim that such domain-adaptive pretraining (DAPT) improves performance on downstream medical tasks, such as answering medical licensing exam questions. In this paper, we compare seven public "medical" LLMs and two VLMs against their corresponding base models, arriving at a different conclusion: all medical VLMs and nearly all medical LLMs fail to consistently improve over their base models in the zero-/few-shot prompting regime for medical question-answering (QA) tasks. For instance, across the tasks and model pairs we consider in the 3-shot setting, medical LLMs only outperform their base models in 12.1% of cases, reach a (statistical) tie in 49.8% of cases, and are significantly worse than their base models in the remaining 38.2% of cases. Our conclusions are based on (i) comparing each medical model head-to-head, directly against the corresponding base model; (ii) optimizing the prompts for each model separately; and (iii) accounting for statistical uncertainty in comparisons. While these basic practices are not consistently adopted in the literature, our ablations show that they substantially impact conclusions. Our findings suggest that state-of-the-art general-domain models may already exhibit strong medical knowledge and reasoning capabilities, and offer recommendations to strengthen the conclusions of future studies.

大模型评估医疗AI提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。