arXiv:2506.14136cs.CV2025-06被引 1

分析BiomedCLIP在放射影像中的表现,发现其在少数类上过预测且区分度差。

Interpreting Biomedical VLMs on High-Imbalance Out-of-Distributions: An Insight into BiomedCLIP on Radiology

  • 通过零样本、微调与线性探测三类设置评估模型表现
  • 零样本下所有标签过预测,精确率低且类别分离差
  • 可视化热图显示模型理解与放射科医生标注有差异

本文构建两个研究目标:一是探索开源大模型BiomedCLIP的嵌入空间,分析其类别可分性;二是量化该模型在高度不平衡、分布外的多标签医学数据集上的局限性。实验基于具有上述特性的IU-xray数据集,评估BiomedCLIP在三种场景下的图像分类能力:零样本推理、全量微调和线性探测。结果表明,零样本设置下模型对所有标签均出现过预测,导致精度低且类间分离性差;全量微调提升了特定疾病识别能力,而线性探测则捕捉到重叠特征。通过Grad-CAM热图进行可视化,并与15份放射科医生标注对比,揭示模型理解偏差。研究强调需谨慎调整模型以提升真实场景下的可靠性与适用性。相关代码已开源并维护于GitHub。

原文摘要 · Abstract (English)

In this paper, we construct two research objectives: i) explore the learned embedding space of BiomedCLIP, an open-source large vision language model, to analyse meaningful class separations, and ii) quantify the limitations of BiomedCLIP when applied to a highly imbalanced, out-of-distribution multi-label medical dataset. We experiment on IU-xray dataset, which exhibits the aforementioned criteria, and evaluate BiomedCLIP in classifying images (radiographs) in three contexts: zero-shot inference, full finetuning, and linear probing. The results show that the model under zero-shot settings over-predicts all labels, leading to poor precision and inter-class separability. Full fine-tuning improves classification of distinct diseases, while linear probing detects overlapping features. We demonstrate visual understanding of the model using Grad-CAM heatmaps and compare with 15 annotations by a radiologist. We highlight the need for careful adaptations of the models to foster reliability and applicability in a real-world setting. The code for the experiments in this work is available and maintained on GitHub.

医学视觉大模型分析零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。