arXiv:2608.25810cs.CV2026-08中稿 · MIRASOL Workshop, …

无需标签即可选出最适合医疗影像分类的模型。

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

论文配图:Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy
图 1 · 摘自论文原文
  • 基于伪标签不一致性构建无标签选择准则,无需目标域标注或微调。
  • 在三种医院间分布偏移场景下,排序相关性高达0.943(p<0.05)。
  • 特别适合标注数据少的资源受限场景,优于传统源域准确率排名。

基础模型在医疗影像分析中应用日益广泛,但在跨机构分布偏移场景下,其性能差异大且无法在无目标域标签时评估。面对多个候选模型和源域有标签数据,如何选择部署于无标签目标域的模型仍无解。本文提出一种基于SUDO框架的无标签选择准则,该框架通过预测概率划分未标记目标数据,并在各区域测量反映类别混淆的伪标签不一致性;综合各区域结果得到无需目标标注或微调的评分(AURCC)。实验表明,AURCC可在零样本与MLP探测两种设置下,对多种视觉-语言模型(BioMedCLIP、CXR-CLIP、CheXzero、MedCLIP、MedImageInsight、CLIP)在胸片分类任务中进行有效排序,在三个跨医院分布偏移场景中,与真实排序的相关性达0.943(p<0.05)。相比基于源域保留精度的基准方法,当源域数据充足时表现相当,而在数据稀少时更具优势,适用于资源受限环境。

原文摘要 · Abstract (English)

Foundation models are increasingly deployed for medical image analysis. However, under the inter-institutional distribution shift typical of deployment, their performance varies widely and cannot be known without target-domain labels, which are rarely available. This leaves a practical question unresolved: given several candidate foundational models and labeled-data from a source domain, which one to deploy in an unlabeled target domain? We propose a label-free selection criterion built on SUDO, a framework for evaluating clinical AI systems without ground-truth annotations. SUDO partitions the unlabeled target data by predicted probability and, for each region, measures a pseudo-label discrepancy reflecting class contamination; aggregated across regions, this yields a score (AURCC) requiring neither target annotation nor fine-tuning. We show that AURCC can be used to rank a variety of vision-language models (BioMedCLIP, CXR-CLIP, CheXzero, MedCLIP, MedImageInsight, CLIP) on chest X-ray classification across three inter-hospital shift scenarios, under zero-shot and MLP-probe regimes. The AURCC ranking recovers the ground-truth ranking with Spearman rho up to 0.943 (p<0.05). Against the natural baseline of ranking by held-out source accuracy, AURCC is competitive when the labeled source is large and yields a more accurate ranking once it is small; the regime of interest in resource-constrained settings.

医疗影像模型选择无标签分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。