arXiv:2505.02654cs.CV2025-05被引 1

让仿真内镜图像更逼真,同时保持解剖结构不变,无需真实标注即可训练分割模型。

Sim2Real in endoscopy segmentation with a novel structure aware image translation

  • 设计新图像转换模型,保留仿真图结构同时添加真实纹理
  • 在结肠镜折叠分割任务中,无需真实标注即实现高精度分割
  • 适用于无公开基准的内镜结构分割研究,适合医学影像领域开发者

内镜图像中解剖标志物的自动分割可辅助医生诊断、治疗或培训。然而,获取真实图像标注费时费力;尽管合成数据标注较易,但模型在真实数据上泛化能力差。现有生成方法虽能增加真实纹理,却难以保持原始场景结构。本文提出一种新型结构感知图像翻译模型,在保持关键场景布局的同时为模拟内镜图像添加真实纹理。该方法生成多种内镜场景下的逼真图像,并成功用于训练无需真实标注的下游分割模型。以结肠镜折叠分割为例,折叠是可能遮挡黏膜和息肉的关键解剖特征。实验在自建仿真数据集与真实数据集EndoMapper(EM)上进行,生成图像有效保留了原始折叠形状与位置,优于现有方法。所有新生成数据及新增EM元数据将公开,以推动该任务的研究,因目前尚无公开基准。

原文摘要 · Abstract (English)

Automatic segmentation of anatomical landmarks in endoscopic images can provide assistance to doctors and surgeons for diagnosis, treatments or medical training. However, obtaining the annotations required to train commonly used supervised learning methods is a tedious and difficult task, in particular for real images. While ground truth annotations are easier to obtain for synthetic data, models trained on such data often do not generalize well to real data. Generative approaches can add realistic texture to it, but face difficulties to maintain the structure of the original scene. The main contribution in this work is a novel image translation model that adds realistic texture to simulated endoscopic images while keeping the key scene layout information. Our approach produces realistic images in different endoscopy scenarios. We demonstrate these images can effectively be used to successfully train a model for a challenging end task without any real labeled data. In particular, we demonstrate our approach for the task of fold segmentation in colonoscopy images. Folds are key anatomical landmarks that can occlude parts of the colon mucosa and possible polyps. Our approach generates realistic images maintaining the shape and location of the original folds, after the image-style-translation, better than existing methods. We run experiments both on a novel simulated dataset for fold segmentation, and real data from the EndoMapper (EM) dataset. All our new generated data and new EM metadata is being released to facilitate further research, as no public benchmark is currently available for the task of fold segmentation.

医学图像图像翻译仿真到真实结肠镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。