arXiv:2505.22079cs.CV2025-05CVPR被引 11

改进CLIP模型,让医疗影像理解更懂否定句和不平衡数据。

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis

  • 用动态软标签增强医学语义表达
  • 引入否定硬负例提升临床语言理解能力
  • 适配多任务场景,尤其适合医疗图像分析

大规模图文对数据集的发展推动了视觉-语言预训练的进展。然而,直接将通用领域模型如CLIP应用于医疗数据时,面临否定句处理与数据分布不均衡的挑战。为此,我们提出结合临床增强的动态软标签与医学图对齐机制,提升对比损失在医疗场景中的适用性;同时引入基于否定的硬负例,深化模型对临床语言复杂性的理解。该方法可无缝集成至医疗CLIP训练流程,在零样本、微调分类及报告检索等多项任务中达到领先性能。为全面评估模型对临床语言的理解能力,我们构建了专用于胸部X光(CXR)的基准数据集CXRA-Align,用于检验否定与临床信息的解析能力。实验表明,所提方法实现简单、泛化性强,有效提升医疗视觉-语言模型的能力,推动医学影像中临床语言理解的发展。

原文摘要 · Abstract (English)

The development of large-scale image-text pair datasets has significantly advanced self-supervised learning in Vision-Language Processing (VLP). However, directly applying general-domain architectures such as CLIP to medical data presents challenges, particularly in handling negations and addressing the inherent data imbalance of medical datasets. To address these issues, we propose a novel approach that integrates clinically-enhanced dynamic soft labels and medical graphical alignment, thereby improving clinical comprehension and the applicability of contrastive loss in medical contexts. Furthermore, we introduce negation-based hard negatives to deepen the model's understanding of the complexities of clinical language. Our approach is easily integrated into the medical CLIP training pipeline and achieves state-of-the-art performance across multiple tasks, including zero-shot, fine-tuned classification, and report retrieval. To comprehensively evaluate our model's capacity for understanding clinical language, we introduce CXR-Align, a benchmark uniquely designed to evaluate the understanding of negation and clinical information within chest X-ray (CXR) datasets. Experimental results demonstrate that our proposed methods are straightforward to implement and generalize effectively across contrastive learning frameworks, enhancing medical VLP capabilities and advancing clinical language understanding in medical imaging.

医疗视觉对比学习临床理解否定识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。