arXiv:2603.06971cs.CV2026-03

让单目内窥镜视频实现精准持续的手术场景三维重建

SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation

  • 用公开立体数据生成伪真实深度图,解决手术数据稀缺问题
  • 混合监督+自校正机制,提升对数据缺陷的鲁棒性
  • 分层推理框架有效抑制长视频中的姿态漂移,适合临床应用

从单目内窥镜视频重建手术场景对推动机器人辅助手术至关重要。然而,现有通用重建模型面临两大挑战:缺乏标注训练数据,以及在长视频序列中性能下降。为此,我们提出SurgCUT3R,一个适配手术领域的系统性框架。首先,利用公开立体手术数据集构建大规模、度量尺度的伪真值深度图,缓解数据短缺。其次,提出混合监督策略,结合伪真值与几何自校正,增强对数据缺陷的鲁棒性。第三,引入分层推理框架,通过两个专用模型分别保障全局稳定性和局部精度,有效缓解长视频中的累积姿态漂移。在SCARED和StereoMIS数据集上的实验表明,该方法在准确率与效率间达到良好平衡,姿态估计接近顶尖水平但速度显著更快,为手术环境下的鲁棒重建提供了实用可行的解决方案。

原文摘要 · Abstract (English)

Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic-assisted surgery. However, the application of state-of-the-art general-purpose reconstruction models is constrained by two key challenges: the lack of supervised training data and performance degradation over long video sequences. To overcome these limitations, we propose SurgCUT3R, a systematic framework that adapts unified 3D reconstruction models to the surgical domain. Our contributions are threefold. First, we develop a data generation pipeline that exploits public stereo surgical datasets to produce large-scale, metric-scale pseudo-ground-truth depth maps, effectively bridging the data gap. Second, we propose a hybrid supervision strategy that couples our pseudo-ground-truth with geometric self-correction to enhance robustness against inherent data imperfections. Third, we introduce a hierarchical inference framework that employs two specialized models to effectively mitigate accumulated pose drift over long surgical videos: one for global stability and one for local accuracy. Experiments on the SCARED and StereoMIS datasets demonstrate that our method achieves a competitive balance between accuracy and efficiency, delivering near state-of-the-art but substantially faster pose estimation and offering a practical and effective solution for robust reconstruction in surgical environments. Project page: https://chumo-xu.github.io/SurgCUT3R-ICRA26/.

三维重建手术影像视觉定位深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。