通过双视角文本提示融合,提升腰椎细粒度分割精度
Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
- 用解剖感知文本提示生成器,将图像标注转为多视图解剖语义提示
- 在SPIDER数据集上达79.39%的Dice,比SOTA方法高8.31个百分点
- 适合医学影像分析、精准诊断场景,尤其关注腰椎结构细分
准确的腰椎分割对脊柱疾病诊断至关重要。现有方法通常采用粗粒度分割策略,缺乏精细细节;且依赖纯视觉模型,难以捕捉解剖语义,导致类别误判和分割不精细。为此,我们提出ATM-Net,一种基于解剖感知、文本引导的多模态融合框架,用于腰椎亚结构(椎体VBs、椎间盘IDs、椎管SC)的细粒度分割。ATM-Net采用解剖感知文本提示生成器(ATPG),自适应地将图像标注转化为不同视角的解剖感知提示。这些提示通过整体解剖语义融合模块(HASF)与图像特征融合,构建完整的解剖上下文。通道级对比解剖增强模块(CCAE)通过类别级通道层面的多模态对比学习,进一步提升类别区分能力并优化分割结果。在MRSpineSeg和SPIDER数据集上的大量实验表明,ATM-Net显著优于现有最优方法,尤其在类别区分和分割细节方面表现突出。例如,在SPIDER数据集上,其Dice达到79.39%,HD95为9.91像素,分别优于竞争方法SpineParseNet 8.31%和4.14像素。
原文摘要 · Abstract (English)
Accurate lumbar spine segmentation is crucial for diagnosing spinal disorders. Existing methods typically use coarse-grained segmentation strategies that lack the fine detail needed for precise diagnosis. Additionally, their reliance on visual-only models hinders the capture of anatomical semantics, leading to misclassified categories and poor segmentation details. To address these limitations, we present ATM-Net, an innovative framework that employs an anatomy-aware, text-guided, multi-modal fusion mechanism for fine-grained segmentation of lumbar substructures, i.e., vertebrae (VBs), intervertebral discs (IDs), and spinal canal (SC). ATM-Net adopts the Anatomy-aware Text Prompt Generator (ATPG) to adaptively convert image annotations into anatomy-aware prompts in different views. These insights are further integrated with image features via the Holistic Anatomy-aware Semantic Fusion (HASF) module, building a comprehensive anatomical context. The Channel-wise Contrastive Anatomy-Aware Enhancement (CCAE) module further enhances class discrimination and refines segmentation through class-wise channel-level multi-modal contrastive learning. Extensive experiments on the MRSpineSeg and SPIDER datasets demonstrate that ATM-Net significantly outperforms state-of-the-art methods, with consistent improvements regarding class discrimination and segmentation details. For example, ATM-Net achieves Dice of 79.39% and HD95 of 9.91 pixels on SPIDER, outperforming the competitive SpineParseNet by 8.31% and 4.14 pixels, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。