解决医学图文预训练中的假负例问题,提升细粒度对齐效果。
FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- 基于文本相似度自适应挖掘正样本,减少假负例。
- 引入文本引导的稀疏注意力池化,实现局部视觉与文本精准对齐。
- 适合医疗影像理解、多模态模型优化的研究者参考。
医学视觉-语言预训练(VLP)通过利用图像与报告配对数据,在推动医学图像理解方面具有巨大潜力。然而,现有方法受限于语义相近文本引发的假负例(FaNe)以及细粒度跨模态对齐不足。为此,我们提出FaNe——一种语义增强型VLP框架。为缓解假负例问题,提出基于文本-文本相似度的语义感知正样本挖掘策略,并采用自适应归一化。进一步设计了文本条件稀疏注意力池化模块,通过文本线索引导的局部视觉表示实现细粒度图像-文本对齐。为增强模态内区分能力,开发了一种硬负样本感知对比损失,自适应重加权语义相似的负样本。在五个下游医学影像基准上的实验表明,FaNe在图像分类、目标检测和语义分割任务上均达到当前最优性能,验证了该框架的有效性。
原文摘要 · Abstract (English)
Medical vision-language pre-training (VLP) offers significant potential for advancing medical image understanding by leveraging paired image-report data. However, existing methods are limited by Fa}lse Negatives (FaNe) induced by semantically similar texts and insufficient fine-grained cross-modal alignment. To address these limitations, we propose FaNe, a semantic-enhanced VLP framework. To mitigate false negatives, we introduce a semantic-aware positive pair mining strategy based on text-text similarity with adaptive normalization. Furthermore, we design a text-conditioned sparse attention pooling module to enable fine-grained image-text alignment through localized visual representations guided by textual cues. To strengthen intra-modal discrimination, we develop a hard-negative aware contrastive loss that adaptively reweights semantically similar negatives. Extensive experiments on five downstream medical imaging benchmarks demonstrate that FaNe achieves state-of-the-art performance across image classification, object detection, and semantic segmentation, validating the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。