让模型聚焦病灶区域,提升乳腺钼靶良恶性分类准确率
Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification

- 用病灶区域与临床描述对齐,学习局部特征
- 在两个数据集上均显著优于现有方法
- 适合需要精准定位和良恶性判断的医学影像分析
基于对比学习的视觉-语言模型在医学图像分析中表现良好,但传统全局图像-文本对齐不适用于乳腺钼靶:诊断关键病灶空间位置明确,仅占图像小部分,整体图像特征会稀释细微形态学线索。本文提出一种检测引导的区域-锚定视觉-语言学习方法,模拟放射科医生诊断流程。首先,在区域-文本对比预训练阶段,将病灶特征与源自放射科元数据的结构化临床描述对齐;为缓解低词汇量下的语义坍塌和背景偏差,引入多组件目标函数,包含正向对齐、细粒度语义硬负样本和背景抑制。其次,联合优化辅助病灶检测头与对比分类任务,保持空间敏感性,实现定位感知的良恶性分类。在独立数据集CBIS-DDSM和VinDr-Mammo上的大量实验表明,该方法在同域、跨数据集及迁移学习场景下均优于现有方法。
原文摘要 · Abstract (English)
Vision-language models trained with contrastive objectives have shown promise in medical image analysis. However, conventional global image-text alignment is ill-suited for mammography, where diagnostically relevant lesions are spatially localized and occupy only a small fraction of the image. Subtle morphological cues critical for malignancy assessment can be diluted when representations are learned at the whole-image level. In this work, we propose a novel region-grounded vision-language learning method for detection-guided mammographic lesion classification. The method mirrors radiologists' diagnostic paradigm. First, a region-text contrastive pretraining stage aligns lesion-specific features with structured clinical descriptors derived from radiology metadata. To mitigate semantic collapse and background bias in low-vocabulary settings, we introduce a multi-component objective incorporating positive alignment, fine-grained semantic hard negatives, and background suppression. Second, an auxiliary lesion detection head is jointly optimized with contrastive classification to preserve spatial sensitivity and enable localization-aware malignancy classification. Extensive experiments on two independent datasets, CBIS-DDSM and VinDr-Mammo, show superior performance of our method compared to related methods under in-domain, cross-dataset, and transfer learning settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。