arXiv:2411.12198cs.CVcs.AI2024-11被引 8

用扩散模型生成可控的结肠镜图像,提升数据多样性与临床真实性

CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis

  • 基于扩散模型设计可控生成框架,支持位置、形状与临床特征精确调控
  • 合成图像在空间一致性与临床描述匹配度上显著优于现有方法
  • 构建多模态数据集,适用于医学影像生成与辅助诊断研究

结肠镜检查对发现腺瘤性息肉、预防结直肠癌至关重要。然而,现有结肠镜数据集规模有限且难以获取,制约了息肉检测模型的鲁棒性发展。尽管已有研究尝试合成结肠镜图像,但当前方法存在生成不稳定、数据多样性不足的问题,且缺乏对生成过程的精确控制,导致图像难以满足临床质量要求。为此,本文提出CCIS-Diff,一种基于扩散架构的可控结肠镜图像生成模型。该方法可精确控制息肉的空间属性(如位置、形状)及临床特征,符合临床描述。我们引入模糊掩码加权策略,实现合成息肉与结肠黏膜的自然融合;同时采用文本感知注意力机制,引导生成结果反映真实临床特征。为支持此目标,我们构建了一个新的多模态结肠镜数据集,包含图像、掩码标注和对应的临床文本描述。实验表明,本方法能生成高质量、多样化的结肠镜图像,在空间约束与临床一致性方面表现优异,为后续分割与诊断任务提供有力支持。

原文摘要 · Abstract (English)

Colonoscopy is crucial for identifying adenomatous polyps and preventing colorectal cancer. However, developing robust models for polyp detection is challenging by the limited size and accessibility of existing colonoscopy datasets. While previous efforts have attempted to synthesize colonoscopy images, current methods suffer from instability and insufficient data diversity. Moreover, these approaches lack precise control over the generation process, resulting in images that fail to meet clinical quality standards. To address these challenges, we propose CCIS-DIFF, a Controlled generative model for high-quality Colonoscopy Image Synthesis based on a Diffusion architecture. Our method offers precise control over both the spatial attributes (polyp location and shape) and clinical characteristics of polyps that align with clinical descriptions. Specifically, we introduce a blur mask weighting strategy to seamlessly blend synthesized polyps with the colonic mucosa, and a text-aware attention mechanism to guide the generated images to reflect clinical characteristics. Notably, to achieve this, we construct a new multi-modal colonoscopy dataset that integrates images, mask annotations, and corresponding clinical text descriptions. Experimental results demonstrate that our method generates high-quality, diverse colonoscopy images with fine control over both spatial constraints and clinical consistency, offering valuable support for downstream segmentation and diagnostic tasks.

图像生成医学影像扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。