arXiv:2501.09008cs.CV2025-01被引 7

SimGen可同时生成高保真手术图像与精确分割图,减少人工标注成本。

SimGen: A Diffusion-Based Framework for Simultaneous Surgical Image and Segmentation Mask Generation

  • 基于扩散模型与残差U-Net,联合生成图像与分割掩码。
  • 在6个公开数据集上优于基线,图像与语义相似度指标更优。
  • 适合需成对手术数据但受限于真实数据的科研与教学场景。

获取和标注外科数据通常耗时耗力、涉及伦理问题且需专家参与。尽管文本到图像等生成式AI能缓解数据稀缺问题,但精准外科应用、仿真与教育仍需包含空间标注(如分割掩码)的数据。本文提出新任务与方法SimGen,实现图像与分割掩码的同时生成。SimGen基于DDPM框架与残差U-Net,利用交叉相关先验捕捉连续图像与离散掩码分布间的依赖关系,并引入标准斐波那契晶格(CFL)提升掩码在RGB空间中的类别可分性与均匀性。实验表明,SimGen在六组公开数据集上均优于基线,图像与语义判别距离指标更优;消融研究显示CFL显著提升掩码质量与空间分离性。下游实验验证生成图像-掩码对在法规限制人类数据使用的场景下具有可用性。该工作为生成配对的外科图像与复杂标签提供了低成本方案,推动外科AI发展,降低对昂贵人工标注的依赖。

原文摘要 · Abstract (English)

Acquiring and annotating surgical data is often resource-intensive, ethical constraining, and requiring significant expert involvement. While generative AI models like text-to-image can alleviate data scarcity, incorporating spatial annotations, such as segmentation masks, is crucial for precision-driven surgical applications, simulation, and education. This study introduces both a novel task and method, SimGen, for Simultaneous Image and Mask Generation. SimGen is a diffusion model based on the DDPM framework and Residual U-Net, designed to jointly generate high-fidelity surgical images and their corresponding segmentation masks. The model leverages cross-correlation priors to capture dependencies between continuous image and discrete mask distributions. Additionally, a Canonical Fibonacci Lattice (CFL) is employed to enhance class separability and uniformity in the RGB space of the masks. SimGen delivers high-fidelity images and accurate segmentation masks, outperforming baselines across six public datasets assessed on image and semantic inception distance metrics. Ablation study shows that the CFL improves mask quality and spatial separation. Downstream experiments suggest generated image-mask pairs are usable if regulations limit human data release for research. This work offers a cost-effective solution for generating paired surgical images and complex labels, advancing surgical AI development by reducing the need for expensive manual annotations.

扩散模型医学图像生成分割掩码外科AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。