arXiv:2509.10344cs.CVcs.AI2025-09中稿 · MICCAI 2025被引 1

用几何引导对齐多视角乳腺影像,提升医学视觉语言模型性能

GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

  • 基于乳腺影像成像原理,设计几何引导的跨视图对齐机制
  • 在EMBED数据集上预训练,多任务表现优于现有方法
  • 适合医学影像与多视图对齐研究者参考

乳腺筛查是早期发现乳腺癌的重要手段。深度学习可提升解读速度与准确性,但受限于医疗图像数据少及自然图像与医学图像之间的领域差异,基础视觉语言模型(VLM)发展受阻。现有乳腺影像VLM多源自自然图像,忽略乳腺影像特有的多视图关系——放射科医生会同时分析双侧视图以判断同侧对应关系,而当前方法常将视图独立处理或未有效建模跨视图对应,丢失关键几何上下文,导致预测效果不佳。本文提出GLAM:一种用于乳腺影像多视图VLM预训练的几何引导局部对齐方法。通过利用乳腺影像成像过程的先验知识,模型在全局与局部层面联合进行视觉-视觉、视觉-语言对比学习,实现跨视图局部对齐与细粒度特征提取。在目前最大的开放乳腺影像数据集EMBED[14]上预训练后,该模型在多个数据集、不同设置下均超越基线方法。

原文摘要 · Abstract (English)

Mammography screening is an essential tool for early detection of breast cancer. The speed and accuracy of mammography interpretation have the potential to be improved with deep learning methods. However, the development of a foundation visual language model (VLM) is hindered by limited data and domain differences between natural and medical images. Existing mammography VLMs, adapted from natural images, often ignore domain-specific characteristics, such as multi-view relationships in mammography. Unlike radiologists who analyze both views together to process ipsilateral correspondence, current methods treat them as independent images or do not properly model the multi-view correspondence learning, losing critical geometric context and resulting in suboptimal prediction. We propose GLAM: Global and Local Alignment for Multi-view mammography for VLM pretraining using geometry guidance. By leveraging the prior knowledge about the multi-view imaging process of mammograms, our model learns local cross-view alignments and fine-grained local features through joint global and local, visual-visual, and visual-language contrastive learning. Pretrained on EMBED [14], one of the largest open mammography datasets, our model outperforms baselines across multiple datasets under different settings.

医学影像多视图对齐视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。