arXiv:2501.14548cs.CV2025-01ICLR被引 69

通过细粒度对齐影像与报告,提升CT图像疾病诊断能力

Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

  • 将解剖区域与报告描述逐个匹配,进行个体化对比预训练
  • 零样本分类平均AUC达81.3%,优于CLIP和有监督方法12.9%和8.0%
  • 适用于多病种、多解剖部位的临床辅助诊断场景

人工智能在提升放射科医生影像解读效率和准确性方面潜力巨大,但通用模型需大规模数据与详尽标注,这在医疗环境中往往难以实现。近期研究利用放射科报告作为高质量监督信号,基于对比语言图像预训练(CLIP)构建语言引导的影像理解模型。然而,这些方法通常将整幅图像与报告进行对比,忽略了影像区域与报告语句间的局部关联,可能影响模型性能与可迁移性。本文提出一种细粒度视觉-语言模型(fVLM),用于解剖级CT图像理解。具体而言,显式匹配CT图像的解剖区域与报告中对应的描述,并对每个解剖部位分别进行对比预训练。细粒度对齐面临大量假阴性挑战,主要源于解剖层面正常样本众多及相似异常表现。为此,我们识别正常与异常样本的假阴性,并从患者级配对转向疾病感知配对以校准对比学习。我们构建了迄今最大的CT数据集,包含69,086名患者的影像与报告数据,并在15个主要解剖部位上评估了54项重要疾病的诊断任务。实验结果表明,fVLM在多任务医学图像理解中具有显著潜力。在零样本分类任务中,54项诊断任务平均AUC达81.3%,较CLIP和有监督方法分别提升12.9%和8.0%。

原文摘要 · Abstract (English)

Artificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively.

CT诊断视觉语言模型细粒度对齐零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。