arXiv:2604.11197cs.CV2026-04中稿 · Medical Image Anal…

让医学影像模型精准理解病灶区域,支持点、框、掩码多种提示。

MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration

论文配图:MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration
图 1 · 摘自论文原文
  • 在特征层融合区域提示,灵活响应不同形式的定位输入。
  • 在640万医学图像上预训练,9730万区域标注实现细粒度理解。
  • 适用于零样本识别与交互式分割,可直接接入大语言模型。

对比语言-图像预训练(CLIP)在全局图像理解与零样本迁移方面表现优异,但医学图像分析的核心在于对特定解剖结构或病灶区域的细粒度理解。因此,准确捕捉医生或感知模型提供的感兴趣区域(RoI)信息至关重要。为此,我们提出MedP-CLIP,一种区域感知的医学视觉语言模型(VLM)。该模型创新性地融合医学先验知识,并设计了特征级区域提示集成机制,可在关注局部区域时保持全局上下文意识,同时灵活应对点、边界框、掩码等多种提示形式。我们在一个精心构建的大规模数据集上进行预训练,该数据集包含超过640万张医学图像和9730万条区域级标注,使模型具备跨疾病、跨模态的细粒度空间语义理解能力。实验表明,MedP-CLIP在零样本识别、交互式分割以及赋能多模态大语言模型等任务中显著优于基线方法。该模型为医学AI提供了一种可扩展、即插即用的视觉主干,结合整体图像理解与精确区域分析。

原文摘要 · Abstract (English)

Contrastive Language-Image Pre-training (CLIP) has demonstrated outstanding performance in global image understanding and zero-shot transfer through large-scale text-image alignment. However, the core of medical image analysis often lies in the fine-grained understanding of specific anatomical structures or lesion regions. Therefore, precisely comprehending region-of-interest (RoI) information provided by medical professionals or perception models becomes crucial. To address this need, we propose MedP-CLIP, a region-aware medical vision-language model (VLM). MedP-CLIP innovatively integrates medical prior knowledge and designs a feature-level region prompt integration mechanism, enabling it to flexibly respond to various prompt forms (e.g., points, bounding boxes, masks) while maintaining global contextual awareness when focusing on local regions. We pre-train the model on a meticulously constructed large-scale dataset (containing over 6.4 million medical images and 97.3 million region-level annotations), equipping it with cross-disease and cross-modality fine-grained spatial semantic understanding capabilities. Experiments demonstrate that MedP-CLIP significantly outperforms baseline methods in various medical tasks, including zero-shot recognition, interactive segmentation, and empowering multimodal large language models. This model provides a scalable, plug-and-play visual backbone for medical AI, combining holistic image understanding with precise regional analysis.

医学视觉区域感知多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。