首个医学多模态上下文学习评测基准,揭示模型在医疗场景中的关键缺陷。
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning
- 由11位医学专家设计111个真实医疗任务,含多模态示例与问答对
- 多数大模型在上下文学习中仅提升8%-9.4%,且易受无关示例干扰
- 发现模型存在显著新近偏差,最后展示的示例影响最大
尽管多模态上下文学习(ICL)在医学等领域的潜力巨大,但研究仍不充分。临床医生常需基于少量案例或有限鉴别诊断进行适应性判断。虽然多模态大语言模型(MLLMs)在医学视觉问答(VQA)上表现进步,其从上下文中学习多模态任务的能力尚不明确。我们提出SMMILE,首个由医学专家驱动的多模态医疗ICL基准。11位医学专家构建了111个问题,涵盖6个专科和13种影像模态,共517个问题-图像-答案三元组。我们进一步推出增强版SMMILE++,包含1038个排列变体。对15个MLLM的全面评估显示,多数模型在医疗任务中表现出中等至较差的多模态ICL能力。在开放评测中,上下文学习仅带来8%(SMMILE)和9.4%(SMMILE++)的平均性能提升。我们发现模型对无关示例高度敏感:单个噪声或无关示例可导致性能下降高达9.5%。此外,模型存在新近偏差——将最相关示例置于最后可使性能提升最高达71%。这些发现揭示了当前MLLM在从上下文学习多模态医疗任务时的关键局限与偏差。SMMILE已开源:https://smmile-benchmark.github.io。
原文摘要 · Abstract (English)
Multimodal in-context learning (ICL) remains underexplored despite significant potential for domains such as medicine. Clinicians routinely encounter diverse, specialized tasks requiring adaptation from limited examples, such as drawing insights from a few relevant prior cases or considering a constrained set of differential diagnoses. While multimodal large language models (MLLMs) have shown advances in medical visual question answering (VQA), their ability to learn multimodal tasks from context is largely unknown. We introduce SMMILE, the first expert-driven multimodal ICL benchmark for medical tasks. Eleven medical experts curated problems, each including a multimodal query and multimodal in-context examples as task demonstrations. SMMILE encompasses 111 problems (517 question-image-answer triplets) covering 6 medical specialties and 13 imaging modalities. We further introduce SMMILE++, an augmented variant with 1038 permuted problems. A comprehensive evaluation of 15 MLLMs demonstrates that most models exhibit moderate to poor multimodal ICL ability in medical tasks. In open-ended evaluations, ICL contributes only an 8% average improvement over zero-shot on SMMILE and 9.4% on SMMILE++. We observe a susceptibility for irrelevant in-context examples: even a single noisy or irrelevant example can degrade performance by up to 9.5%. Moreover, we observe that MLLMs are affected by a recency bias, where placing the most relevant example last can lead to substantial performance improvements of up to 71%. Our findings highlight critical limitations and biases in current MLLMs when learning multimodal medical tasks from context. SMMILE is available at https://smmile-benchmark.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。