arXiv:2409.03412cs.CVphysics.med-ph2024-09被引 1

用器官描述文本提升医学图像分割准确率

TG-LMM: Enhancing Medical Image Segmentation Accuracy through Text-Guided Large Multi-Modal Model

  • 引入器官空间位置的文本描述作为先验知识
  • 在三个数据集上优于MedSAM、SAM等模型
  • 适合需要高精度分割的临床医生和研究者

我们提出TG-LMM(Text-Guided Large Multi-Modal Model),一种利用器官文本描述提升医学图像分割准确率的新方法。现有自动分割模型未能有效利用器官位置等先验知识;此前的图文模型侧重目标识别而非分割精度提升;已有模型尝试融合先验知识但未结合预训练模型。TG-LMM将专家提供的器官空间位置描述融入分割流程,采用预训练图像与文本编码器以减少训练参数并加速训练,并设计了全面的图像-文本信息融合结构,实现双模态深度整合。我们在三个权威医学图像数据集上评估该模型,涵盖人体多部位分割任务,结果表明其性能显著优于MedSAM、SAM和nnUnet等现有方法。

原文摘要 · Abstract (English)

We propose TG-LMM (Text-Guided Large Multi-Modal Model), a novel approach that leverages textual descriptions of organs to enhance segmentation accuracy in medical images. Existing medical image segmentation methods face several challenges: current medical automatic segmentation models do not effectively utilize prior knowledge, such as descriptions of organ locations; previous text-visual models focus on identifying the target rather than improving the segmentation accuracy; prior models attempt to use prior knowledge to enhance accuracy but do not incorporate pre-trained models. To address these issues, TG-LMM integrates prior knowledge, specifically expert descriptions of the spatial locations of organs, into the segmentation process. Our model utilizes pre-trained image and text encoders to reduce the number of training parameters and accelerate the training process. Additionally, we designed a comprehensive image-text information fusion structure to ensure thorough integration of the two modalities of data. We evaluated TG-LMM on three authoritative medical image datasets, encompassing the segmentation of various parts of the human body. Our method demonstrated superior performance compared to existing approaches, such as MedSAM, SAM and nnUnet.

医学图像分割多模态文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。