arXiv:2602.23203cs.CVcs.AI2026-02被引 2

用扩散模型生成逼真结肠镜视频,解决数据少、动态不连贯问题。

ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation

  • 分帧建模+内容感知,精准控制病变特征与动态变化。
  • 生成视频过渡自然,跨数据集验证表现优于现有方法。
  • 适合医学影像生成、临床数据增强研究者使用。

结肠镜视频生成可提供动态、信息丰富的数据,对肠道疾病诊断至关重要,尤其在数据稀缺场景下。高质量视频生成需兼顾时间一致性与临床属性的精确控制,但受限于肠道结构不规则、病灶表现多样及成像模态差异。为此,我们提出ColoDiff,一种基于扩散模型的框架,实现动态一致且内容感知的结肠镜视频生成,以缓解数据短缺并辅助临床分析。在帧间层面,TimeStream模块通过跨帧标记化机制解耦时间依赖,实现对不规则肠道结构的复杂动态建模;在帧内层面,Content-Aware模块引入噪声注入嵌入与可学习原型,实现对临床属性的精细控制,突破扩散模型粗粒度引导的局限。此外,ColoDiff采用非马尔可夫采样策略,采样步骤减少超90%,支持实时生成。在三个公开数据集和一个医院数据库上评估,涵盖生成指标与下游任务(疾病诊断、模态区分、肠道准备评分、病灶分割)。大量实验表明,ColoDiff生成视频具有平滑过渡与丰富动态。本工作推动可控结肠镜视频生成,揭示合成视频在补充真实数据与缓解临床数据稀缺方面的潜力。

原文摘要 · Abstract (English)

Colonoscopy video generation delivers dynamic, information-rich data critical for diagnosing intestinal diseases, particularly in data-scarce scenarios. High-quality video generation demands temporal consistency and precise control over clinical attributes, but faces challenges from irregular intestinal structures, diverse disease representations, and various imaging modalities. To this end, we propose ColoDiff, a diffusion-based framework that generates dynamic-consistent and content-aware colonoscopy videos, aiming to alleviate data shortage and assist clinical analysis. At the inter-frame level, our TimeStream module decouples temporal dependency from video sequences through a cross-frame tokenization mechanism, enabling intricate dynamic modeling despite irregular intestinal structures. At the intra-frame level, our Content-Aware module incorporates noise-injected embeddings and learnable prototypes to realize precise control over clinical attributes, breaking through the coarse guidance of diffusion models. Additionally, ColoDiff employs a non-Markovian sampling strategy that cuts steps by over 90% for real-time generation. ColoDiff is evaluated across three public datasets and one hospital database, based on both generation metrics and downstream tasks including disease diagnosis, modality discrimination, bowel preparation scoring, and lesion segmentation. Extensive experiments show ColoDiff generates videos with smooth transitions and rich dynamics. ColoDiff presents an effort in controllable colonoscopy video generation, revealing the potential of synthetic videos in complementing authentic representation and mitigating data scarcity in clinical settings.

视频生成医学影像扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。