用无标签数据提升医学图像少样本适配效果,减少一半以上标注工作量。
Semi-Supervised Few-Shot Adaptation of Vision-Language Models
- 通过文本引导的伪标签传播,实现半监督少样本适配
- 在极低样本下使模型性能显著提升,标注成本降低50%以上
- 特别适合标注成本高的医疗图像任务
视觉-语言模型(VLMs)在大规模异构数据上预训练后,可生成丰富的多模态嵌入,有效迁移至新任务。在医学影像领域,专用VLMs已在零样本和少样本图像分类中表现优异,有助于缓解专家标注的高成本问题。然而,在极端少样本场景下,类别分布不均仍导致少数类表现不佳,影响整体性能。为此,我们提出一种高效的半监督求解器,利用无标签数据在少样本适配过程中传播文本引导的伪标签。该方法显著降低标注需求,在低样本条件下将标注工作量减少超过50%。
原文摘要 · Abstract (English)
Vision-language models (VLMs) pre-trained on large, heterogeneous data sources are becoming increasingly popular, providing rich multi-modal embeddings that enable efficient transfer to new tasks. A particularly relevant application is few-shot adaptation, where only a handful of annotated examples are available to adapt the model through multi-modal linear probes. In medical imaging, specialized VLMs have shown promising performance in zero- and few-shot image classification, which is valuable for mitigating the high cost of expert annotations. However, challenges remain in extremely low-shot regimes: the inherent class imbalances in medical tasks often lead to underrepresented categories, penalizing overall model performance. To address this limitation, we propose leveraging unlabeled data by introducing an efficient semi-supervised solver that propagates text-informed pseudo-labels during few-shot adaptation. The proposed method enables lower-budget annotation pipelines for adapting VLMs, reducing labeling effort by >50% in low-shot regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。