用结构生成图像,让肠镜深度估计更准。
Structure-to-Image: Zero-Shot Depth Estimation in Colonoscopy via High-Fidelity Sim-to-Real Adaptation
- 把深度图当生成基础,而非仅作约束。
- 零样本测试下误差降低44.18%。
- 适合做医学影像深度估计的研究者。
肠镜单目深度估计受仿真与真实图像间域差距困扰。现有图像到图像转换方法以深度为后验约束,常因难以平衡真实感与结构一致性,导致结构失真和反光伪影。为此,我们提出结构生成图像范式,将深度图从被动约束转为主动生成基础。首次在肠镜领域引入相位一致性,并设计跨层级结构约束,协同优化几何结构与血管纹理等细粒度细节。在公开的体模数据集上进行零样本评估,经我们生成数据微调的深度模型,相比竞品最大可降低44.18%的均方根误差(RMSE)。代码已开源:https://github.com/YyangJJuan/PC-S2I.git。
原文摘要 · Abstract (English)
Monocular depth estimation (MDE) for colonoscopy is hampered by the domain gap between simulated and real-world images. Existing image-to-image translation methods, which use depth as a posterior constraint, often produce structural distortions and specular highlights by failing to balance realism with structure consistency. To address this, we propose a Structure-to-Image paradigm that transforms the depth map from a passive constraint into an active generative foundation. We are the first to introduce phase congruency to colonoscopic domain adaptation and design a cross-level structure constraint to co-optimize geometric structures and fine-grained details like vascular textures. In zero-shot evaluations conducted on a publicly available phantom dataset, the MDE model that was fine-tuned on our generated data achieved a maximum reduction of 44.18% in RMSE compared to competing methods. Our code is available at https://github.com/YyangJJuan/PC-S2I.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。