用语义图生成手术场景图像,解决数据少且不平衡问题
Image Synthesis with Class-Aware Semantic Diffusion Models for Surgical Scene Segmentation
- 以分割图为条件生成图像,聚焦关键组织类别
- 新损失函数提升小类图像质量,生成结果更真实
- 支持文本提示生成多类分割图,适合医学图像增强
手术场景分割对提升手术精度至关重要,但常受数据稀缺与不平衡影响。现有基于生成对抗网络和扩散模型的语义图像生成方法往往生成图像多样性不足,难以捕捉细微、关键的组织类别,限制了实际效果。为此,本文提出类感知语义扩散模型(CASDM),利用分割图作为条件进行图像合成,以应对数据稀缺与不平衡问题。设计了类感知均方误差与类感知自感知损失函数,优先关注较少见但重要的组织类别,从而提升图像质量和相关性。此外,首次以新颖方式通过文本提示生成多类分割图以指定内容,再由CASDM生成手术场景图像,用于训练和验证分割模型的数据集增强。评估结果显示,该方法在图像质量与下游分割性能方面均表现出显著有效性与泛化能力,在多样且具有挑战性的数据集上有效推进了手术场景分割技术。
原文摘要 · Abstract (English)
Surgical scene segmentation is essential for enhancing surgical precision, yet it is frequently compromised by the scarcity and imbalance of available data. To address these challenges, semantic image synthesis methods based on generative adversarial networks and diffusion models have been developed. However, these models often yield non-diverse images and fail to capture small, critical tissue classes, limiting their effectiveness. In response, we propose the Class-Aware Semantic Diffusion Model (CASDM), a novel approach which utilizes segmentation maps as conditions for image synthesis to tackle data scarcity and imbalance. Novel class-aware mean squared error and class-aware self-perceptual loss functions have been defined to prioritize critical, less visible classes, thereby enhancing image quality and relevance. Furthermore, to our knowledge, we are the first to generate multi-class segmentation maps using text prompts in a novel fashion to specify their contents. These maps are then used by CASDM to generate surgical scene images, enhancing datasets for training and validating segmentation models. Our evaluation, which assesses both image quality and downstream segmentation performance, demonstrates the strong effectiveness and generalisability of CASDM in producing realistic image-map pairs, significantly advancing surgical scene segmentation across diverse and challenging datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。