调准阈值后,零样本医学视觉语言模型能超出现有轻量CNN。
Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis
- 用验证集优化分类阈值,激活零样本模型潜力。
- 肺炎检测F1达0.8841,超轻量CNN的0.8803。
- 适合想用零样本模型做医疗诊断的研究者。
利用自动化方法准确解读胸部X光片是医学影像中的关键任务。本文对比了监督式轻量级卷积神经网络(CNN)与前沿零样本医学视觉语言模型BiomedCLIP在两项诊断任务中的表现:在PneumoniaMNIST数据集上检测肺炎,在Shenzhen TB数据集上检测结核病。实验表明,监督式CNN在两种情况下均表现强劲。尽管默认零样本性能较低,但通过在验证集上优化分类阈值,BiomedCLIP性能显著提升。在肺炎检测中,校准后零样本模型F1分数达0.8841,优于监督式CNN的0.8803;在结核病检测中,F1从0.4812提升至0.7684,接近监督基线的0.7834。研究揭示:恰当校准对释放零样本视觉语言模型的诊断能力至关重要,使其可媲美甚至超越高效的任务专用监督模型。
原文摘要 · Abstract (English)
The accurate interpretation of chest radiographs using automated methods is a critical task in medical imaging. This paper presents a comparative analysis between a supervised lightweight Convolutional Neural Network (CNN) and a state-of-the-art, zero-shot medical Vision-Language Model (VLM), BiomedCLIP, across two distinct diagnostic tasks: pneumonia detection on the PneumoniaMNIST benchmark and tuberculosis detection on the Shenzhen TB dataset. Our experiments show that supervised CNNs serve as highly competitive baselines in both cases. While the default zero-shot performance of the VLM is lower, we demonstrate that its potential can be unlocked via a simple yet crucial remedy: decision threshold calibration. By optimizing the classification threshold on a validation set, the performance of BiomedCLIP is significantly boosted across both datasets. For pneumonia detection, calibration enables the zero-shot VLM to achieve a superior F1-score of 0.8841, surpassing the supervised CNN's 0.8803. For tuberculosis detection, calibration dramatically improves the F1-score from 0.4812 to 0.7684, bringing it close to the supervised baseline's 0.7834. This work highlights a key insight: proper calibration is essential for leveraging the full diagnostic power of zero-shot VLMs, enabling them to match or even outperform efficient, task-specific supervised models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。