用文本提示提升医学图像分割精度,解决语义理解不足问题
Text-guided multi-stage cross-perception network for medical image segmentation
- 分阶段跨模态注意力机制增强细粒度语义理解
- 多阶段对齐损失使图文特征在不同层级保持一致
- 在三个数据集上超越现有方法,最高达88.09%分割准确率
医学图像分割在临床诊断、治疗规划和疾病监测中至关重要。传统方法如U-Net因目标区域语义表达能力弱、泛化性差且缺乏交互性而受限。引入文本提示可更精准定位病灶,但现有方法仍存在跨模态交互不足、特征表示不充分的问题。为此,我们提出文本引导的多阶段跨感知网络(TMC),通过多阶段跨注意力模块(MCM)增强模型对细粒度语义的理解,并设计多阶段对齐损失(MA Loss)提升不同特征层级间跨模态语义的一致性。在三个公开数据集(QaTa-COV19、MosMedData、Duke-Breast-Cancer-MRI)上的实验表明,TMC表现优异,分别取得84.65%、78.39%、88.09%的Dice分数,持续优于基于U-Net的网络及现有文本引导方法。
原文摘要 · Abstract (English)
Medical image segmentation plays a crucial role in clinical medicine, serving as a key tool for auxiliary diagnosis, treatment planning, and disease monitoring. However, traditional segmentation methods such as U-Net are often limited by weak semantic expression of target regions, which stems from insufficient generalization and a lack of interactivity. Incorporating text prompts offers a promising avenue to more accurately pinpoint lesion locations, yet existing text-guided methods are still hindered by insufficient cross-modal interaction and inadequate cross-modal feature representation. To address these challenges, we propose the Text-guided Multi-stage Cross-perception network (TMC). TMC incorporates a Multi-stage Cross-attention Module (MCM) to enhance the model's understanding of fine-grained semantic details and a Multi-stage Alignment Loss (MA Loss) to improve the consistency of cross-modal semantics across different feature levels. Experimental results on three public datasets (QaTa-COV19, MosMedData, and Duke-Breast-Cancer-MRI) demonstrate the superior performance of TMC, achieving Dice scores of 84.65\%, 78.39\%, and 88.09\%, respectively, and consistently outperforming both U-Net-based networks and existing text-guided methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。