通过优化样本与锚点对齐,提升医学影像零样本分类准确率
ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification

- 用精选负样本构建病灶专注数据集,减少标签共现偏差
- 在新数据集上对多种病灶的AUC提升2-10个百分点,气胸达0.94
- 无需大规模重训练,适合临床部署的稳定零样本迁移
零样本视觉语言模型(VLM)在胸部X光分类中展现潜力,但受限于标签共现混淆、长尾类别不平衡及域偏移下的迁移不稳定性。本文提出ProtoCLIP,一种针对CLIP类VLM的精炼策略,通过目标数据筛选与蒸馏锚点对齐提升零样本判别能力。具体地,构建聚焦病灶的训练子集,加入精心设计的负样本以降低共现偏差;同时引入保持表示结构的蒸馏目标,稳定适配过程并增强对临床相关共现病灶的区分能力。在未见过的VinDr-CXR数据集上评估,相较于强基线模型,ProtoCLIP在多个病灶上的AUC提升2-10个百分点。尤其对于气胸,达到0.94的当前最优AUC。结果表明,锚点引导的精炼结合精选监督与受控适配,可在不需大规模重训练的前提下缓解医疗VLM的常见零样本迁移失败问题。
原文摘要 · Abstract (English)
Zero-shot vision-language models (VLMs) have shown promise for chest radiograph classification, but their performance is often limited by confounding label co-occurrence, long-tail class imbalance, and transfer instability under domain shift. We propose ProtoCLIP, a refinement strategy for CLIP-style VLMs that improves zero-shot discrimination through targeted data curation and distilled anchor alignment. Specifically, we construct pathology-focused training subsets with curated negative samples to reduce co-occurrence bias. We also introduce a representation-preserving distillation objective to stabilize adaptation while maintaining semantic structure and improving discrimination of clinically relevant co-occurring pathologies. Evaluated on an unseen dataset VinDr-CXR, ProtoCLIP improves AUC by 2-10 percentage points over a strong CLIP-based baseline across multiple findings. For pneumothorax specifically, ProtoCLIP achieves a state-of-the-art AUC of 0.94. These results demonstrate that anchor-guided refinement, coupled with curated supervision and controlled adaptation, can mitigate common zero-shot transfer failures in medical VLMs without requiring large-scale retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。