arXiv:2412.13533cs.CV2024-12被引 6

通过多层级对比对齐,让文本更精准指导医学图像分割

Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments

  • 引入目标感知的语义距离模块,细化图文对齐
  • 多尺度对比对齐使文本能指导图像细节区域
  • 适合需要精准语义引导的医学图像分割场景

医学图像分割是众多医学工程应用的基础任务。近年来,基于语言的分割在临床报告易得的医疗场景中展现出潜力,报告中的诊断信息可作为语义引导。然而,现有方法忽视了图像与文本模态间的内在模式差异,导致视觉-语言融合效果不佳。对比学习虽能对齐图文模式,但当前技术多聚焦高层全局语义,缺乏对关键病灶等局部细节的细粒度指导。为此,本文提出目标感知多层级对比对齐框架(TMCA),通过三项创新实现:(i) 目标敏感的语义距离模块,利用病灶信息实现更精细的图文对齐;(ii) 多层级对比对齐策略,将文本指导精确到多尺度图像细节;(iii) 语言引导的目标增强模块,基于对齐结果强化对关键区域的关注。在四个公开基准数据集上的大量实验表明,TMCA在性能上显著优于现有最先进方法。

原文摘要 · Abstract (English)

Medical image segmentation is a fundamental task in numerous medical engineering applications. Recently, language-guided segmentation has shown promise in medical scenarios where textual clinical reports are readily available as semantic guidance. Clinical reports contain diagnostic information provided by clinicians, which can provide auxiliary textual semantics to guide segmentation. However, existing language-guided segmentation methods neglect the inherent pattern gaps between image and text modalities, resulting in sub-optimal visual-language integration. Contrastive learning is a well-recognized approach to align image-text patterns, but it has not been optimized for bridging the pattern gaps in medical language-guided segmentation that relies primarily on medical image details to characterize the underlying disease/targets. Current contrastive alignment techniques typically align high-level global semantics without involving low-level localized target information, and thus cannot deliver fine-grained textual guidance on crucial image details. In this study, we propose a Target-informed Multi-level Contrastive Alignment framework (TMCA) to bridge image-text pattern gaps for medical language-guided segmentation. TMCA enables target-informed image-text alignments and fine-grained textual guidance by introducing: (i) a target-sensitive semantic distance module that utilizes target information for more granular image-text alignment modeling, (ii) a multi-level contrastive alignment strategy that directs fine-grained textual guidance to multi-scale image details, and (iii) a language-guided target enhancement module that reinforces attention to critical image regions based on the aligned image-text patterns. Extensive experiments on four public benchmark datasets demonstrate that TMCA enabled superior performance over state-of-the-art language-guided medical image segmentation methods.

医学图像语言引导对比学习分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。