arXiv:2503.23534cs.CVcs.AI2025-03被引 2

用双向渐进融合提升医学影像分割中的图文对齐效果

BiPVL-Seg: Bidirectional Progressive Vision-Language Fusion with Global-Local Alignment for Medical Image Segmentation

  • 双向渐进融合让视觉与文本特征分阶段交互
  • 全局局部对比对齐使文本编码更理解医学概念
  • 适合需要图文联合分析的医学影像研究者

医学图像分割通常仅依赖视觉数据,忽略了临床诊断中丰富的文本信息。尽管视觉语言模型试图弥合这一差距,但现有方法常独立处理视觉与文本特征,导致跨模态对齐弱。简单融合策略因空间视觉特征与序列文本嵌入的本质差异而失效。此外,医学术语不同于通用语言,限制了现成文本编码器的效果,进一步阻碍图文对齐。我们提出BiPVL-Seg,一个端到端框架,通过架构与训练创新实现视觉语言融合与嵌入对齐,两者相互增强以提升医学图像分割性能。该框架引入双向渐进融合,促进视觉与文本编码器间分阶段信息交换;同时采用全局-局部对比对齐训练目标,通过在类别与概念层级上对齐文本与视觉嵌入,增强文本编码器的理解能力。在涵盖CT与MR模态的多个医学影像基准上的大量实验表明,相比当前最优方法,BiPVL-Seg在复杂多类别分割任务中表现更优。源代码已公开于GitHub仓库。

原文摘要 · Abstract (English)

Medical image segmentation typically relies solely on visual data, overlooking the rich textual information clinicians use for diagnosis. Vision-language models attempt to bridge this gap, but existing approaches often process visual and textual features independently, resulting in weak cross-modal alignment. Simple fusion techniques fail due to the inherent differences between spatial visual features and sequential text embeddings. Additionally, medical terminology deviates from general language, limiting the effectiveness of off-the-shelf text encoders and further hindering vision-language alignment. We propose BiPVL-Seg, an end-to-end framework that integrates vision-language fusion and embedding alignment through architectural and training innovations, where both components reinforce each other to enhance medical image segmentation. BiPVL-Seg introduces bidirectional progressive fusion in the architecture, which facilitates stage-wise information exchange between vision and text encoders. Additionally, it incorporates global-local contrastive alignment, a training objective that enhances the text encoder's comprehension by aligning text and vision embeddings at both class and concept levels. Extensive experiments on diverse medical imaging benchmarks across CT and MR modalities demonstrate BiPVL-Seg's superior performance when compared with state-of-the-art methods in complex multi-class segmentation. Source code is available in this GitHub repository.

医学图像图文对齐分割视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。