无需标定或先验,实时重建内窥镜视频的3D结构
Endo3R: Unified Online Reconstruction from Dynamic Monocular Endoscopic Video
- 用双记忆机制实现长期动态场景的增量式重建
- 在SCARED和Hamlyn数据集上实现零样本深度预测与位姿估计
- 适合需要实时3D手术辅助的临床场景
从单目手术视频重建3D场景可增强医生感知,在多种计算机辅助手术任务中至关重要。然而,由于内窥镜视频存在动态形变和无纹理表面等固有问题,实现尺度一致的重建仍是挑战。现有方法依赖标定或器械先验,或采用类似SfM的多阶段流程,导致误差累积且需离线优化。本文提出Endo3R,一种统一的在线3D基础模型,无需任何先验或额外优化,即可实现尺度一致的单目手术视频重建。该模型通过全局对齐点图、尺度一致的视频深度与相机参数联合预测完成重建。核心贡献在于将近期成对重建模型扩展至长期增量式动态重建,引入基于不确定性的双记忆机制,分别维护短期动态与长期空间一致性历史令牌。针对手术场景高度动态性,利用Sampson距离衡量令牌不确定性并剔除高不确定性项。鉴于缺乏带真值深度与相机位姿的内窥镜数据集,设计了一种新型动态感知光流损失的自监督机制。在SCARED与Hamlyn数据集上的大量实验表明,该方法在零样本手术视频深度预测与相机位姿估计上表现优越,同时具备在线效率。
原文摘要 · Abstract (English)
Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open challenge due to inherent issues in endoscopic videos, such as dynamic deformations and textureless surfaces. Despite recent advances, current methods either rely on calibration or instrument priors to estimate scale, or employ SfM-like multi-stage pipelines, leading to error accumulation and requiring offline optimization. In this paper, we present Endo3R, a unified 3D foundation model for online scale-consistent reconstruction from monocular surgical video, without any priors or extra optimization. Our model unifies the tasks by predicting globally aligned pointmaps, scale-consistent video depths, and camera parameters without any offline optimization. The core contribution of our method is expanding the capability of the recent pairwise reconstruction model to long-term incremental dynamic reconstruction by an uncertainty-aware dual memory mechanism. The mechanism maintains history tokens of both short-term dynamics and long-term spatial consistency. Notably, to tackle the highly dynamic nature of surgical scenes, we measure the uncertainty of tokens via Sampson distance and filter out tokens with high uncertainty. Regarding the scarcity of endoscopic datasets with ground-truth depth and camera poses, we further devise a self-supervised mechanism with a novel dynamics-aware flow loss. Abundant experiments on SCARED and Hamlyn datasets demonstrate our superior performance in zero-shot surgical video depth prediction and camera pose estimation with online efficiency. Project page: https://wrld.github.io/Endo3R/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。