arXiv:2504.16047cs.CVcs.AI2025-04被引 1

对比三款医学视觉语言模型,发现预训练方式决定其在影像诊断中的表现优劣。

Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis

  • 用自监督或文本监督方式预训练的模型分别擅长分割与分类任务。
  • 无文本监督的RAD-DINO在气胸分割上表现最佳,精度显著提升。
  • 融合全局与局部特征的定制模型可普遍提升所有模型性能,尤其对复杂分割任务。

基础模型通过大规模数据和自监督方法训练,已成为推动医学AI发展的前沿方向。本研究评估了三种视觉语言基础模型(RAD-DINO、CheXagent、BiomedCLIP)在胸部X光片中识别气胸和心脏扩大的细粒度影像特征的能力,涵盖分类、分割和回归任务。自监督的RAD-DINO在分割任务中表现最优,而文本监督的CheXagent在分类任务中更优,BiomedCLIP则表现不一致。通过引入融合全局与局部特征的定制分割模型,所有基础模型性能均得到提升,尤其在具有挑战性的气胸分割任务中效果显著。结果表明,预训练方法显著影响模型在特定下游任务上的表现:无文本监督模型更适合细粒度分割,而文本监督模型在分类和可解释性方面更具优势。这些发现为临床放射学中根据具体应用选择合适模型提供了指导。

原文摘要 · Abstract (English)

Foundation models, trained on vast amounts of data using self-supervised techniques, have emerged as a promising frontier for advancing artificial intelligence (AI) applications in medicine. This study evaluates three different vision-language foundation models (RAD-DINO, CheXagent, and BiomedCLIP) on their ability to capture fine-grained imaging features for radiology tasks. The models were assessed across classification, segmentation, and regression tasks for pneumothorax and cardiomegaly on chest radiographs. Self-supervised RAD-DINO consistently excelled in segmentation tasks, while text-supervised CheXagent demonstrated superior classification performance. BiomedCLIP showed inconsistent performance across tasks. A custom segmentation model that integrates global and local features substantially improved performance for all foundation models, particularly for challenging pneumothorax segmentation. The findings highlight that pre-training methodology significantly influences model performance on specific downstream tasks. For fine-grained segmentation tasks, models trained without text supervision performed better, while text-supervised models offered advantages in classification and interpretability. These insights provide guidance for selecting foundation models based on specific clinical applications in radiology.

医学影像视觉语言模型分割预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。