用文字指令精准分割3D医学影像,轻量高效还懂语义。
SwinTF3D: A Lightweight Multimodal Fusion Approach for Text-Guided 3D Medical Image Segmentation
- 视觉与文本双模态融合,通过轻量结构对齐语言提示与解剖位置。
- 在BTCV数据集上多器官分割Dice超90%,效率优于传统3D Transformer。
- 适合需要灵活交互、计算资源有限的临床场景使用。
人工智能在医学影像中的应用推动了自动化器官分割的显著进展。然而,现有大多数3D分割框架仅依赖大规模标注数据的视觉学习,限制了其在新领域和临床任务中的适应性。模型缺乏语义理解能力,难以应对用户自定义的分割目标。为此,我们提出SwinTF3D,一种轻量级多模态融合方法,统一视觉与语言表征,实现文本引导的3D医学图像分割。该模型采用基于Transformer的视觉编码器提取体数据特征,并通过高效的融合机制将特征与紧凑的文本编码器结合。此设计使系统能够理解自然语言提示,准确对齐语义线索与医学体数据中的空间结构,同时以低计算开销生成精确、上下文感知的分割结果。在BTCV数据集上的大量实验表明,尽管架构紧凑,SwinTF3D在多个器官上仍取得具有竞争力的Dice和IoU分数。模型对未见数据泛化能力强,相比传统Transformer分割网络具备显著效率优势。通过融合视觉感知与语言理解,SwinTF3D建立了一种可交互、可解释的文本驱动3D医学图像分割范式,为临床影像中更灵活、高效解决方案开辟了新路径。
原文摘要 · Abstract (English)
The recent integration of artificial intelligence into medical imaging has driven remarkable advances in automated organ segmentation. However, most existing 3D segmentation frameworks rely exclusively on visual learning from large annotated datasets restricting their adaptability to new domains and clinical tasks. The lack of semantic understanding in these models makes them ineffective in addressing flexible, user-defined segmentation objectives. To overcome these limitations, we propose SwinTF3D, a lightweight multimodal fusion approach that unifies visual and linguistic representations for text-guided 3D medical image segmentation. The model employs a transformer-based visual encoder to extract volumetric features and integrates them with a compact text encoder via an efficient fusion mechanism. This design allows the system to understand natural-language prompts and correctly align semantic cues with their corresponding spatial structures in medical volumes, while producing accurate, context-aware segmentation results with low computational overhead. Extensive experiments on the BTCV dataset demonstrate that SwinTF3D achieves competitive Dice and IoU scores across multiple organs, despite its compact architecture. The model generalizes well to unseen data and offers significant efficiency gains compared to conventional transformer-based segmentation networks. Bridging visual perception with linguistic understanding, SwinTF3D establishes a practical and interpretable paradigm for interactive, text-driven 3D medical image segmentation, opening perspectives for more adaptive and resource-efficient solutions in clinical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。