arXiv:2512.21135cs.CVcs.AI2025-12

用文本指导医学图像分割,提升准确率且参数更少。

TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation

  • 结合视觉与文本特征,通过结构-语义协同编码增强细节
  • 在5个数据集上达到顶尖性能,关键指标Dice显著提升
  • 适合医疗影像分析、多模态模型研究者参考

文本引导的医学图像分割通过利用临床报告作为辅助信息来提升分割精度。然而,现有方法通常依赖未对齐的图像和文本编码器,需复杂交互模块实现多模态融合。尽管CLIP提供了预对齐的多模态特征空间,但其直接应用于医学影像存在三大局限:细粒度解剖结构保留不足、复杂临床描述建模能力弱、领域特定语义错位。为此,我们提出TGC-Net,一个基于CLIP的结构感知、语义对齐的轻量级框架。具体包括:(1)语义-结构协同编码器(SSE),在CLIP ViT基础上增加CNN分支以实现多尺度结构细化;(2)领域增强文本编码器(DATE),注入大语言模型生成的医学知识;(3)视觉-语言校准模块(VLCM),在统一特征空间中优化跨模态对应关系。在胸片和胸部CT跨模态的五个数据集上的实验表明,TGC-Net在显著减少可训练参数的同时达到当前最优性能,尤其在挑战性基准上取得显著的Dice系数提升。

原文摘要 · Abstract (English)

Text-guided medical segmentation enhances segmentation accuracy by utilizing clinical reports as auxiliary information. However, existing methods typically rely on unaligned image and text encoders, which necessitate complex interaction modules for multimodal fusion. While CLIP provides a pre-aligned multimodal feature space, its direct application to medical imaging is limited by three main issues: insufficient preservation of fine-grained anatomical structures, inadequate modeling of complex clinical descriptions, and domain-specific semantic misalignment. To tackle these challenges, we propose TGC-Net, a CLIP-based framework focusing on parameter-efficient, task-specific adaptations. Specifically, it incorporates a Semantic-Structural Synergy Encoder (SSE) that augments CLIP's ViT with a CNN branch for multi-scale structural refinement, a Domain-Augmented Text Encoder (DATE) that injects large-language-model-derived medical knowledge, and a Vision-Language Calibration Module (VLCM) that refines cross-modal correspondence in a unified feature space. Experiments on five datasets across chest X-ray and thoracic CT modalities demonstrate that TGC-Net achieves state-of-the-art performance with substantially fewer trainable parameters, including notable Dice gains on challenging benchmarks.

医学图像分割文本引导多模态轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。