arXiv:2506.17500cs.CV2025-06被引 5

提出无需平衡数据和验证集的医疗视觉语言模型少样本适配新方法。

Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation

  • 不依赖平衡支持集与验证集,模拟真实医疗数据分布。
  • 现有方法在真实场景下性能下降,有时甚至不如零样本推理。
  • 提出无训练线性探测器,自适应融合视觉与文本监督,高效鲁棒。

视觉语言模型(VLMs)在医学图像分析中日益受到关注。这些模型在大规模异构数据上预训练,生成丰富且可迁移的表征。特别是模态专用VLMs结合少样本适配,已实现高效高精度解决方案。然而,现有研究假设适配数据分布理想化:一是要求平衡的支持集,违背现实疾病发病率的自然不平衡;二是依赖额外验证集以确定关键超参数,导致数据效率低下。本文挑战这些理想化部署条件,提出一种真实、不平衡且无验证的适配设置。跨多种模态与下游任务的广泛基准测试表明,当前方法在真实条件下系统性退化,有时甚至劣于零样本推理。为此,我们引入一种无需训练的线性探测器,能自适应融合视觉与文本监督。详细实验证明该方法是强而高效的基线,可在挑战性场景中实现稳健适配。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are gaining attention in medical image analysis. These are pre-trained on large, heterogeneous data sources, yielding rich and transferable representations. Notably, the combination of modality-specialized VLMs with few-shot adaptation has provided fruitful results, enabling the efficient deployment of high-performing solutions. However, previous works on this topic make strong assumptions about the distribution of adaptation data, which are unrealistic in the medical domain. First, prior art assumes access to a balanced support set, a condition that breaks the natural imbalance in disease prevalence found in real-world scenarios. Second, these works typically assume the presence of an additional validation set to fix critical hyper-parameters, which is highly data-inefficient. This work challenges these favorable deployment scenarios and introduces a realistic, imbalanced, validation-free adaptation setting. Our extensive benchmark across various modalities and downstream tasks demonstrates that current methods systematically compromise their performance when operating under realistic conditions, occasionally even performing worse than zero-shot inference. Also, we introduce a training-free linear probe that adaptively blends visual and textual supervision. Detailed studies demonstrate that the proposed solver is a strong, efficient baseline, enabling robust adaptation in challenging scenarios.

视觉语言模型少样本学习医疗影像无验证适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。