arXiv:2608.13690cs.CVcs.AI2026-08中稿 · BMVC-2026

让医学影像分割融合临床文本,提升精准度。

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

论文配图:MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
图 1 · 摘自论文原文
  • 视觉与语言双向融合,全程协同学习
  • 在多种医学影像上达到领先性能
  • 适合需要临床知识的医疗AI研究者

医学图像分割长期被视为纯视觉问题,但临床诊断依赖解剖、位置、形态及上下文等文本知识。现有基于视觉-语言模型的方法多将语言作为后期条件信号,限制了其对视觉表示的影响。本文提出MedPlex(医学视觉语言协同网络),一个端到端的视觉-语言模型框架,使语言指导成为分割学习中持续且临床相关的组成部分。通过双向融合(Bi-Fusion),视觉与文本表征在整个编码层次中共同演化。MedPlex还引入类别级与区域级概念对齐机制,在不同粒度上组织共享表征:类别级对齐将每个解剖目标锚定至综合临床概念特征,区域级对齐则通过类别特定的视觉证据保留形状、位置、外观和纹理等个体概念。由此,语言在编码器中提供结构化监督,而非仅作为后期提示。MedPlex在多器官、心脏亚结构及肿瘤分割的CT和MR基准测试中均取得最优表现,包括真实自由文本临床监督场景。代码已开源。

原文摘要 · Abstract (English)

Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.

医学分割视觉语言模型临床知识融合双向融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。