arXiv:2607.16327cs.CV2026-07

让文本描述中的位置信息直接指导医学图像分割,提升精准度。

Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation

论文配图:Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation
图 1 · 摘自论文原文
  • 通过多尺度定位任务显式提取文本中的目标位置信息
  • 三种分层融合策略使分割结果更贴合解剖位置
  • 在三个数据集上超越现有方法,适合临床辅助诊断场景

医学图像分割对现代计算机辅助诊疗至关重要。近年来,文本引导分割通过引入临床医生撰写的文本报告作为语义指引,展现出巨大潜力。这些报告包含目标形态、位置及邻近解剖结构的描述,为定位与勾画提供明确指导。现有方法通常通过预训练文本编码器隐式提取语义,并采用简单的图像-文本特征融合,未能显式捕捉文本中嵌入的目标导向信息,尤其缺乏对目标位置的关注,且仅使用基础特征级融合,限制了关键语义的提取与整合。为此,本文提出LoG框架,一种融合定位信息的视觉-语言语义融合方法。该框架通过联合执行多尺度目标定位任务,显式捕获目标导向的视-语义信息,并实现三级定位增强的语义融合:(i) 定位引导的特征融合,将位置相关语义注入视觉特征;(ii) 定位门控注意力融合,利用多尺度定位预测强化关键区域;(iii) 定位约束损失融合,基于空间一致性监督分割结果。在三个主流基准数据集上的实验表明,覆盖三种医学影像模态并配有文本报告,LoG始终优于当前最优方法。

原文摘要 · Abstract (English)

Medical image segmentation is essential for modern computer-aided medicine. Recently, text-guided segmentation has shown promise by incorporating clinician-formulated textual reports as semantic guidance for image segmentation. These textual reports contain language descriptions about the appearance, location, and neighboring anatomy of segmentation targets, providing explicit guidance for target localization and delineation. Existing text-guided segmentation methods typically extract textual semantics implicitly through a pretrained text encoder and then integrate vision-language semantics via straightforward image-text feature fusion. However, these methods do not explicitly capture target-oriented information embedded in textual reports, particularly target location, and do not explore multi-level information fusion strategies beyond basic feature-level fusion, limiting the extraction and integration of critical textual semantics. In this study, we propose LoG, a localization-infused vision-language fusion framework for text-guided medical image segmentation. By jointly performing multi-scale target localization tasks, LoG explicitly captures target-oriented vision-language semantics and enables three-level localization-infused semantic fusion: (i) localization-guided feature fusion that directly infuses location-relevant semantics into visual features, (ii) localization-gated attention fusion that redirects multi-scale localization predictions to reinforce critical regions, and (iii) localization-constrained loss fusion that supervises segmentation based on spatial consistency with target localization. Extensive experiments on three well-established benchmark datasets, involving three medical imaging modalities with paired textual reports, demonstrate that LoG consistently outperforms state-of-the-art medical image segmentation methods.

医学图像文本引导语义融合定位增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。