arXiv:2512.22981cs.CV2025-12

解决医学影像分割中文本定位不准的问题,提升对空间关系的识别能力。

Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation

  • 采用对称最优传输机制建立图像区域与多语义文本的双向对应
  • 引入复合方向引导策略,通过区域级掩码显式约束空间位置
  • 在包含空间关系的病变分割任务中达到当前最优效果

文本引导的医学图像分割利用丰富的临床文本弥补数据稀缺问题,展现出巨大潜力。然而现有方法存在两大瓶颈:一方面难以同时处理诊断性与描述性文本,导致难以准确定位病灶并建立与图像区域的关联;另一方面忽略空间约束,造成严重误判,例如“左下肺”可能被错误覆盖双侧肺部。为此,本文提出空间感知对称对齐(SSA)框架,增强对包含位置、描述和诊断信息的混合文本的建模能力。具体而言,设计对称最优传输对齐机制,强化图像区域与多个相关表达之间的双向细粒度跨模态对应;同时提出复合方向引导策略,通过构建区域级引导掩码显式引入文本中的空间约束。在公开基准上的大量实验表明,SSA在具有空间关系约束的病灶分割任务中达到当前最优性能。

原文摘要 · Abstract (English)

Text-guided Medical Image Segmentation has shown considerable promise for medical image segmentation, with rich clinical text serving as an effective supplement for scarce data. However, current methods have two key bottlenecks. On one hand, they struggle to process diagnostic and descriptive texts simultaneously, making it difficult to identify lesions and establish associations with image regions. On the other hand, existing approaches focus on lesions description and fail to capture positional constraints, leading to critical deviations. Specifically, with the text "in the left lower lung", the segmentation results may incorrectly cover both sides of the lung. To address the limitations, we propose the Spatial-aware Symmetric Alignment (SSA) framework to enhance the capacity of referring hybrid medical texts consisting of locational, descriptive, and diagnostic information. Specifically, we propose symmetric optimal transport alignment mechanism to strengthen the associations between image regions and multiple relevant expressions, which establishes bi-directional fine-grained multimodal correspondences. In addition, we devise a composite directional guidance strategy that explicitly introduces spatial constraints in the text by constructing region-level guidance masks. Extensive experiments on public benchmarks demonstrate that SSA achieves state-of-the-art (SOTA) performance, particularly in accurately segmenting lesions characterized by spatial relational constraints.

医学图像分割文本引导空间约束跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。