让医学影像理解更精准,关注病灶区域而非整体图像。
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
- 引入区域感知的对比学习框架,融合局部病灶与全局语义信息。
- 在5个基准任务上超越现有模型,零样本分类准确率提升显著。
- 适合医疗AI研究者及需要细粒度医学图像分析的临床场景。
医学图像理解对自动化诊断和数据驱动的临床决策支持至关重要。然而,其进展受到两大挑战制约:高质量标注医学数据稀缺,以及过度依赖全局图像特征,常忽略细微但具有临床意义的病灶区域。为此,我们提出RegionMed-CLIP,一种区域感知的多模态对比学习预训练框架,显式融合局部病灶信号与整体语义表征。核心是创新的感兴趣区域(ROI)处理器,自适应整合细粒度区域特征与全局上下文,辅以渐进式训练策略,增强层级化的多模态对齐。为支持大规模区域级表征学习,我们构建了包含广泛区域标注和多层次临床描述的MedRegion-500k医学图文语料库。在图像文本检索、零样本分类和视觉问答任务上的大量实验表明,RegionMed-CLIP在多个任务上均显著优于当前最先进模型。结果凸显了区域感知对比预训练的关键价值,并将RegionMed-CLIP定位为推进多模态医学图像理解的坚实基础。
原文摘要 · Abstract (English)
Medical image understanding plays a crucial role in enabling automated diagnosis and data-driven clinical decision support. However, its progress is impeded by two primary challenges: the limited availability of high-quality annotated medical data and an overreliance on global image features, which often miss subtle but clinically significant pathological regions. To address these issues, we introduce RegionMed-CLIP, a region-aware multimodal contrastive learning framework that explicitly incorporates localized pathological signals along with holistic semantic representations. The core of our method is an innovative region-of-interest (ROI) processor that adaptively integrates fine-grained regional features with the global context, supported by a progressive training strategy that enhances hierarchical multimodal alignment. To enable large-scale region-level representation learning, we construct MedRegion-500k, a comprehensive medical image-text corpus that features extensive regional annotations and multilevel clinical descriptions. Extensive experiments on image-text retrieval, zero-shot classification, and visual question answering tasks demonstrate that RegionMed-CLIP consistently exceeds state-of-the-art vision language models by a wide margin. Our results highlight the critical importance of region-aware contrastive pre-training and position RegionMed-CLIP as a robust foundation for advancing multimodal medical image understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。