arXiv:2411.07546cs.CVcs.AI2024-11被引 3

用正负文本提示减少医学图像异常检测的误报

Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection

  • 通过正负文本提示控制视觉注意力,定位病变区域
  • 在BMAD数据集上显著降低假阳性率,提升检测性能
  • 方法简单有效,适合临床医学图像分析场景

预训练的视觉-语言模型如对比图像-文本预训练(CLIP)可通过文本提示完成多种下游任务,如图像检索与区域定位。尽管具备强大的多模态能力,但在医学领域仍受限于正常区域引发的假阳性问题。为此,已有BioMedCLIP、MedCLIP-SAMv2等变体出现,但假阳性现象依然存在。本文提出一种简单的对比语言提示方法(CLAP),利用正负文本提示协同优化注意力机制:通过正提示聚焦潜在病灶区域,同时借助负提示抑制正常区域的注意力,从而减少误报。在BMAD数据集及六个生物医学基准上的大量实验表明,该方法显著提升了异常检测性能。未来工作将探索自动微调提示的方法以增强实用性。

原文摘要 · Abstract (English)

A pre-trained visual-language model, contrastive language-image pre-training (CLIP), successfully accomplishes various downstream tasks with text prompts, such as finding images or localizing regions within the image. Despite CLIP's strong multi-modal data capabilities, it remains limited in specialized environments, such as medical applications. For this purpose, many CLIP variants-i.e., BioMedCLIP, and MedCLIP-SAMv2-have emerged, but false positives related to normal regions persist. Thus, we aim to present a simple yet important goal of reducing false positives in medical anomaly detection. We introduce a Contrastive LAnguage Prompting (CLAP) method that leverages both positive and negative text prompts. This straightforward approach identifies potential lesion regions by visual attention to the positive prompts in the given image. To reduce false positives, we attenuate attention on normal regions using negative prompts. Extensive experiments with the BMAD dataset, including six biomedical benchmarks, demonstrate that CLAP method enhances anomaly detection performance. Our future plans include developing an automated fine prompting method for more practical usage.

医学图像异常检测提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。