用扩散模型提升肠镜视频深度估计一致性,可直接用于临床3D重建。
ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
- 基于扩散模型学习合成数据中的几何先验,实现视频时序一致的深度图生成。
- 在C3VD数据集上零样本性能领先,超越通用与专用方法。
- 适用于真实肠镜视频,支持3D点云生成与表面覆盖评估,临床价值高。
肠镜检查中的三维场景理解面临重大挑战,亟需自动化方法实现精确深度估计。然而,现有内窥镜深度估计模型在视频序列中难以保持时间一致性,限制了其在三维重建中的应用。本文提出ColonCrafter,一种基于扩散模型的深度估计方法,可从单目肠镜视频生成时序一致的深度图。该方法通过合成肠镜序列学习鲁棒的几何先验,并引入风格迁移技术,在保留几何结构的同时将真实临床视频适配至合成训练域。ColonCrafter在C3VD数据集上实现当前最优的零样本性能,优于通用及专用方法。尽管完整轨迹三维重建仍具挑战,我们展示了其在3D点云生成与表面覆盖评估等临床相关任务中的应用价值。
原文摘要 · Abstract (English)
Three-dimensional (3D) scene understanding in colonoscopy presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle with temporal consistency across video sequences, limiting their applicability for 3D reconstruction. We present ColonCrafter, a diffusion-based depth estimation model that generates temporally consistent depth maps from monocular colonoscopy videos. Our approach learns robust geometric priors from synthetic colonoscopy sequences to generate temporally consistent depth maps. We also introduce a style transfer technique that preserves geometric structure while adapting real clinical videos to match our synthetic training domain. ColonCrafter achieves state-of-the-art zero-shot performance on the C3VD dataset, outperforming both general-purpose and endoscopy-specific approaches. Although full trajectory 3D reconstruction remains a challenge, we demonstrate clinically relevant applications of ColonCrafter, including 3D point cloud generation and surface coverage assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。