对比视觉语言模型与传统模型在肠镜息肉检测中的表现
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images
- 用11种模型对比分析2258张肠镜图像,涵盖CNN、ML和VLM
- ResNet50在息肉检测中表现最优(F1 91.35%),GPT-4优于其他通用VLM
- VLM适合无训练条件下的检测任务,但分类能力仍弱于专用CNN
本研究系统评估了视觉语言模型(VLMs)与经典卷积神经网络(CNNs)及机器学习模型(CMLs)在结肠镜息肉计算机辅助检测(CADe)和诊断(CADx)中的表现。基于428名患者的2,258张结肠镜图像及病理报告,采用标准化预处理方法,比较了11种模型:ResNet50、4种传统机器学习模型(随机森林、支持向量机、逻辑回归、决策树)、两种专用对比视觉语言编码器(CLIP、BiomedCLIP)以及三种通用VLM(GPT-4、Gemini-1.5-Pro、Claude-3-Opus)。结果表明,在息肉检测任务中,ResNet50表现最佳(F1: 91.35%,AUROC: 0.98),BiomedCLIP次之(F1: 88.68%);GPT-4表现接近传统方法(F1: 81.02%),优于其他通用VLM。在分类任务中,性能整体下降,但排名一致:ResNet50最优(加权F1: 74.94%),GPT-4中等(加权F1: 41.18%),显著高于Claude-3-Opus(25.54%)和Gemini 1.5 Pro(6.17%)。结论显示,CNN在两项任务中仍占优势,但特定VLM如BiomedCLIP和GPT-4可在无法训练CNN时用于检测。
原文摘要 · Abstract (English)
Introduction: This study provides a comprehensive performance assessment of vision-language models (VLMs) against established convolutional neural networks (CNNs) and classic machine learning models (CMLs) for computer-aided detection (CADe) and computer-aided diagnosis (CADx) of colonoscopy polyp images. Method: We analyzed 2,258 colonoscopy images with corresponding pathology reports from 428 patients. We preprocessed all images using standardized techniques (resizing, normalization, and augmentation) and implemented a rigorous comparative framework evaluating 11 distinct models: ResNet50, 4 CMLs (random forest, support vector machine, logistic regression, decision tree), two specialized contrastive vision language encoders (CLIP, BiomedCLIP), and three general-purpose VLMs ( GPT-4 Gemini-1.5-Pro, Claude-3-Opus). Our performance assessment focused on two clinical tasks: polyp detection (CADe) and classification (CADx). Result: In polyp detection, ResNet50 achieved the best performance (F1: 91.35%, AUROC: 0.98), followed by BiomedCLIP (F1: 88.68%, AUROC: [AS1] ). GPT-4 demonstrated comparable effectiveness to traditional machine learning approaches (F1: 81.02%, AUROC: [AS2] ), outperforming other general-purpose VLMs. For polyp classification, performance rankings remained consistent but with lower overall metrics. ResNet50 maintained the highest efficacy (weighted F1: 74.94%), while GPT-4 demonstrated moderate capability (weighted F1: 41.18%), significantly exceeding other VLMs (Claude-3-Opus weighted F1: 25.54%, Gemini 1.5 Pro weighted F1: 6.17%). Conclusion: CNNs remain superior for both CADx and CADe tasks. However, VLMs like BioMedCLIP and GPT-4 may be useful for polyp detection tasks where training CNNs is not feasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。