arXiv:2512.17121cs.LG2025-12

研究发现CLIP在医学影像中处理否定句能力差,通过微调提升其表现。

The Effect of Negation on CLIP in Medical Imaging: Limitations of Contrastive Language-Image Pretraining

  • 用临床相关提示微调CLIP模型,增强对否定语句的理解
  • 微调后否定句检索准确率提升,正向提示准确率略有下降
  • 结合注意力分析与特征可视化,揭示文本编码器的改进步骤

大型视觉-语言模型如CLIP因其无需大量标注数据即可实现图像与文本对齐的能力,在医学影像任务中日益广泛应用,适用于图像检索、报告生成和分类等临床场景。然而,这类模型在解析否定短语时表现不佳,这在医疗诊断中尤为关键。本研究评估了Stanford AIMI CheXagent模型在使用含否定与不含否定提示时,对胸部X光片的检索能力。目标是识别其失效原因,并基于先前工作中的微调方法改进其检索精度。结果表明,经过微调后,模型对否定语句的处理能力得到提升,同时正向提示的准确性略有下降。此外,通过词元归因、t-SNE投影和注意力头消融分析,进一步揭示了不同微调策略如何重塑文本编码器对临床否定语言的表征。本研究旨在深入理解CLIP内部机制,提升其在临床语境下对否定语言的处理能力,从而增强医学AI设备的可靠性。

原文摘要 · Abstract (English)

Large vision-language models like CLIP are increasingly used in medical imaging tasks due to their ability to align images and text without the need for extensive labeled data. This makes them particularly useful for applications like image retrieval, report generation, and classification in clinical settings. A potential issue to this approach is that CLIP-based models often under perform when interpreting negated phrases, which is especially problematic in the context of medical diagnosing. In this study, we evaluate the Stanford AIMI CheXagent model on its ability to correctly retrieve chest X-ray images using prompts with and without negation. The goal of this project is to understand where this model fails and then use it as a base model to improve its retrieval accuracy by fine tuning methods outlined in previous work. Results from this study show improvement in handling of negation in the CLIP model with a slight decrease in accuracy of positive prompt evaluation. Alongside retrieval accuracy, we examined internal model behavior through token attribution, t-SNE projection, and attention-head ablation to better characterize how each fine tuning approach reshaped the text encoders representation of negated clinical language. Through this work, we hope to better understand the internal behavior of CLIP and improve its handling of negation using clinically relevant language for improving its reliability in medical AI devices.

医学影像CLIP否定理解微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。