arXiv:2412.20418eess.IVcs.CV2024-12被引 5

用扩散模型合成对齐的多模态肝肿瘤图像,解决临床数据不配准难题。

Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment

  • 通过扩散模型生成配准的多模态肝肿瘤图像,无需严格对齐原始数据。
  • 在公开与内部数据集上,分割性能超越现有最先进方法。
  • 适合医学影像分析、肝肿瘤分割及多模态数据增强研究者。

多模态学习因提供不同视角而提升多种临床任务性能,但现有方法依赖严格配准的多模态数据,这在真实临床图像中难以实现,尤其针对边界模糊的肝肿瘤区域。本文提出Diff4MMLiTS,一个四阶段多模态肝肿瘤分割流程:先对多模态CT进行目标器官预配准;再对标注模态的病灶掩码进行膨胀,并用于修复生成无肿瘤的正常多模态CT;接着基于多模态特征和随机生成的肿瘤掩码,利用潜在扩散模型合成严格对齐的含肿瘤多模态CT;最后训练分割模型,从而摆脱对严格对齐数据的依赖。在公开与内部数据集上的大量实验表明,Diff4MMLiTS优于其他最先进多模态分割方法。

原文摘要 · Abstract (English)

Multimodal learning has been demonstrated to enhance performance across various clinical tasks, owing to the diverse perspectives offered by different modalities of data. However, existing multimodal segmentation methods rely on well-registered multimodal data, which is unrealistic for real-world clinical images, particularly for indistinct and diffuse regions such as liver tumors. In this paper, we introduce Diff4MMLiTS, a four-stage multimodal liver tumor segmentation pipeline: pre-registration of the target organs in multimodal CTs; dilation of the annotated modality's mask and followed by its use in inpainting to obtain multimodal normal CTs without tumors; synthesis of strictly aligned multimodal CTs with tumors using the latent diffusion model based on multimodal CT features and randomly generated tumor masks; and finally, training the segmentation model, thus eliminating the need for strictly aligned multimodal data. Extensive experiments on public and internal datasets demonstrate the superiority of Diff4MMLiTS over other state-of-the-art multimodal segmentation methods.

肝肿瘤分割多模态扩散模型图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。