提升医学影像零样本检测准确率,降低误判。
OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment
- 引入非对角项损失增强正常样本聚类
- 在VinDr-CXR上使AUC提升0.61,误报率下降
- 无需修改模型结构,适合临床部署
对比语言-图像预训练(CLIP)已实现放射科零样本分类,减少对手动标注的依赖。然而,传统对比学习因严格的一致性对齐,导致正常样本聚类混乱,产生高误报和漏报。为此,我们提出OFF-CLIP,通过引入非对角项损失提升正常样本聚类效果,并采用句级文本过滤去除异常报告中错配的正常描述,缓解漏报问题。该方法可直接应用于现有放射科CLIP模型,无需架构修改。实验表明,OFF-CLIP在VinDr-CXR数据集上相较CARZero基线实现0.61的AUC提升,同时保持或改善异常分类性能。此外,其在零样本定位任务中提升指认游戏准确率,验证了更好的异常定位能力。结果表明,OFF-CLIP是医疗视觉-语言模型的有效且高效的增强方案。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-Training (CLIP) has enabled zero-shot classification in radiology, reducing reliance on manual annotations. However, conventional contrastive learning struggles with normal case detection due to its strict intra-sample alignment, which disrupts normal sample clustering and leads to high false positives (FPs) and false negatives (FNs). To address these issues, we propose OFF-CLIP, a contrastive learning refinement that improves normal detection by introducing an off-diagonal term loss to enhance normal sample clustering and applying sentence-level text filtering to mitigate FNs by removing misaligned normal statements from abnormal reports. OFF-CLIP can be applied to radiology CLIP models without requiring any architectural modifications. Experimental results show that OFF-CLIP significantly improves normal classification, achieving a 0.61 Area under the curve (AUC) increase on VinDr-CXR over CARZero, the state-of-the-art zero-shot classification baseline, while maintaining or improving abnormal classification performance. Additionally, OFF-CLIP enhances zero-shot grounding by improving pointing game accuracy, confirming better anomaly localization. These results demonstrate OFF-CLIP's effectiveness as a robust and efficient enhancement for medical vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。