首个可条件生成内镜视频的框架,提升诊断辅助效果。
EndoGen: Conditional Autoregressive Endoscopic Video Generation
- 采用时空网格帧建模策略,将多帧生成转为网格化图像生成。
- 在真实内镜数据上生成高质量视频,提升息肉分割任务性能。
- 适合医疗影像生成、智能诊疗系统研发人员使用。
内镜视频生成对推动医学影像发展和提升诊断能力至关重要。然而,以往研究或仅关注静态图像,缺乏实际应用所需的动态信息;或依赖无条件生成,无法为临床医生提供有意义的参考。为此,本文提出首个条件化内镜视频生成框架——EndoGen。具体而言,构建一种基于自回归架构的模型,并引入定制化的时空网格-帧模式(SGP)策略,将多帧生成学习重构为基于网格的图像生成模式,有效利用自回归结构的全局依赖建模能力。此外,提出语义感知标记掩码(SAT)机制,在生成过程中有选择地聚焦于语义有意义区域,增强内容丰富性与多样性。通过大量实验,验证了该框架在生成高质量、条件引导的内镜内容方面的有效性,并提升了下游息肉分割任务的性能。代码已公开于 https://www.github.com/CUHK-AIM-Group/EndoGen。
原文摘要 · Abstract (English)
Endoscopic video generation is crucial for advancing medical imaging and enhancing diagnostic capabilities. However, prior efforts in this field have either focused on static images, lacking the dynamic context required for practical applications, or have relied on unconditional generation that fails to provide meaningful references for clinicians. Therefore, in this paper, we propose the first conditional endoscopic video generation framework, namely EndoGen. Specifically, we build an autoregressive model with a tailored Spatiotemporal Grid-Frame Patterning (SGP) strategy. It reformulates the learning of generating multiple frames as a grid-based image generation pattern, which effectively capitalizes the inherent global dependency modeling capabilities of autoregressive architectures. Furthermore, we propose a Semantic-Aware Token Masking (SAT) mechanism, which enhances the model's ability to produce rich and diverse content by selectively focusing on semantically meaningful regions during the generation process. Through extensive experiments, we demonstrate the effectiveness of our framework in generating high-quality, conditionally guided endoscopic content, and improves the performance of downstream task of polyp segmentation. Code released at https://www.github.com/CUHK-AIM-Group/EndoGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。