arXiv:2504.11825eess.IVcs.CV2025-04被引 2

用文字指导扩散模型分割3D医学影像,提升精度与效率

TextDiffSeg: Text-guided Latent Diffusion Model for 3d Medical Images Segmentation

  • 通过文本引导的扩散框架融合图像与语言信息
  • 在肾癌、胰腺癌等分割任务中超越现有方法
  • 适合需要精准解剖结构识别的临床诊断场景

扩散概率模型(DPMs)在3D医学图像分割中展现出巨大潜力,但其高计算成本及难以充分捕捉全局3D上下文信息限制了实际应用。为此,我们提出一种新型文本引导扩散模型框架TextDiffSeg。该方法采用条件扩散框架,将3D体数据与自然语言描述结合,实现跨模态嵌入,并建立视觉与文本模态间的共享语义空间。通过引入创新的标签嵌入技术和跨模态注意力机制,显著提升模型对复杂解剖结构的识别能力,有效降低计算复杂度同时保持3D上下文完整性。实验结果表明,TextDiffSeg在肾肿瘤、胰腺肿瘤及多器官分割任务中持续优于现有方法。消融研究进一步验证关键组件的有效性,凸显文本融合、图像特征提取器与标签编码器之间的协同作用。TextDiffSeg为3D医学图像分割提供了高效且准确的解决方案,适用于临床诊断与治疗规划。

原文摘要 · Abstract (English)

Diffusion Probabilistic Models (DPMs) have demonstrated significant potential in 3D medical image segmentation tasks. However, their high computational cost and inability to fully capture global 3D contextual information limit their practical applications. To address these challenges, we propose a novel text-guided diffusion model framework, TextDiffSeg. This method leverages a conditional diffusion framework that integrates 3D volumetric data with natural language descriptions, enabling cross-modal embedding and establishing a shared semantic space between visual and textual modalities. By enhancing the model's ability to recognize complex anatomical structures, TextDiffSeg incorporates innovative label embedding techniques and cross-modal attention mechanisms, effectively reducing computational complexity while preserving global 3D contextual integrity. Experimental results demonstrate that TextDiffSeg consistently outperforms existing methods in segmentation tasks involving kidney and pancreas tumors, as well as multi-organ segmentation scenarios. Ablation studies further validate the effectiveness of key components, highlighting the synergistic interaction between text fusion, image feature extractor, and label encoder. TextDiffSeg provides an efficient and accurate solution for 3D medical image segmentation, showcasing its broad applicability in clinical diagnosis and treatment planning.

3D分割扩散模型文本引导医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。