无需训练即可提升医疗视觉语言模型在分布外数据上的表现
TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models

- 基于少量样本修正推理时的类别得分,不引入额外可训练参数
- 在九个医学影像数据集上显著改善分布外性能,多数情况优于有训练的方法
- 适合低数据场景,对不同医学模态和模型都有效
医疗视觉语言模型(VLMs)具备出色的零样本能力,但其在分布外(OOD)数据上的表现仍因领域偏移和预训练带来的类别偏差而下降。现有少样本适应方法通常引入额外可训练组件,在极低数据场景(如1样本)下不稳定,且在不同医学数据上鲁棒性不足。我们提出TCLA,一种完全无需训练的少样本适应方法,具有快速、模型无关的特点。TCLA基于少量支持样本修正推理时的类别得分,通过增强类别区分度并减少领域偏移,提升预训练VLM的性能。在涵盖X光、超声、MRI、CT、组织病理学等多模态的九个数据集上进行的大量实验表明,TCLA能持续提升医疗VLM在分布外数据上的表现,多数情况下优于现有基于训练的适应方法。
原文摘要 · Abstract (English)
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a small set of support samples, boosting pretrained VLMs performance by improving inter-class deconfusion and reducing domain shift. Extensive experiments on nine datasets across multiple medical imaging modalities including X-ray, Ultrasound, MRI, CT, Histopathology, demonstrate that TCLA consistently improves OOD performance of Medical VLMs and, in most of cases, outperforms existing training-based adaptation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。