首个医学视觉语言模型少样本适配基准,验证文本引导线性探针有效性
Few-shot Adaptation of Medical Vision-Language Models
- 提出基于文本嵌入与视觉原型融合的可学习分类加权机制
- 在9个下游任务上超越复杂提示学习和适配器方法,速度更快
- 支持黑盒场景,适用于医疗图像少样本迁移研究
通过多模态学习整合图像与文本数据已成为医学影像研究的新范式,继其在计算机视觉中的成功应用之后。尽管已有大量工作致力于建立医学基础模型并实现零样本迁移,但当前主流的少样本设置仍相对未被探索。借鉴计算机视觉中少样本设定的兴起,本文首次构建了严格少样本条件下医学视觉语言模型(VLMs)适配的结构化基准,并评估了常用于自然图像的多种适配策略。此外,我们评估了一种简单的线性探针基线改进方案,该方案通过可学习的类别专属乘子优化视觉原型与文本嵌入的融合。令人意外的是,这种文本引导的线性探针在性能上媲美复杂的提示学习与适配器方法,且运行速度显著更快,同时支持黑盒场景。我们的大规模实验涵盖三种不同医学模态、多个专用基础模型、九项下游任务及数种先进少样本适配方法。相关基准与代码已开源,以推动该新兴领域的进一步发展: https://github.com/FereshteShakeri/few-shot-MedVLMs。
原文摘要 · Abstract (English)
Integrating image and text data through multi-modal learning has emerged as a new approach in medical imaging research, following its successful deployment in computer vision. While considerable efforts have been dedicated to establishing medical foundation models and their zero-shot transfer to downstream tasks, the popular few-shot setting remains relatively unexplored. Following on from the currently strong emergence of this setting in computer vision, we introduce the first structured benchmark for adapting medical vision-language models (VLMs) in a strict few-shot regime and investigate various adaptation strategies commonly used in the context of natural images. Furthermore, we evaluate a simple generalization of the linear-probe adaptation baseline, which seeks an optimal blending of the visual prototypes and text embeddings via learnable class-wise multipliers. Surprisingly, such a text-informed linear probe yields competitive performances in comparison to convoluted prompt-learning and adapter-based strategies, while running considerably faster and accommodating the black-box setting. Our extensive experiments span three different medical modalities and specialized foundation models, nine downstream tasks, and several state-of-the-art few-shot adaptation methods. We made our benchmark and code publicly available to trigger further developments in this emergent subject: \url{https://github.com/FereshteShakeri/few-shot-MedVLMs}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。