arXiv:2503.07204cs.CV2025-03被引 10

首个同时实现内窥镜深度与位姿估计的自监督学习框架。

Endo-FASt3r: Endoscopic Foundation model Adaptation for Structure from motion

  • 基于基础模型扩展,采用新型高秩适配技术加速收敛。
  • 在SCARED数据集上位姿估计提升10%,深度估计提升2%。
  • 适用于多类内窥镜数据集,对机器人手术3D重建有实用价值。

精准的深度与相机位姿估计对实现机器人辅助手术中的高质量三维可视化至关重要。尽管已有研究通过自监督学习(SSL)将基础模型适配于单目内窥镜场景的深度估计,但尚未有工作探索其在位姿估计中的应用。现有方法依赖低秩适配,限制了模型更新空间。本文提出Endo-FASt3r,首个同时利用基础模型进行单目深度与位姿估计的自监督学习框架。通过改进Reloc3r相对位姿估计基础模型,设计出Reloc3rX以支持在自监督学习下的稳定收敛,并提出新的适配技术DoMoRA,实现更高秩更新与更快收敛。在SCARED数据集上的实验表明,该方法在位姿估计上相比先前工作提升10%,深度估计提升2%。在Hamlyn与StereoMIS数据集上也获得相似性能增益,验证了其跨数据集的泛化能力。

原文摘要 · Abstract (English)

Accurate depth and camera pose estimation is essential for achieving high-quality 3D visualisations in robotic-assisted surgery. Despite recent advancements in foundation model adaptation to monocular depth estimation of endoscopic scenes via self-supervised learning (SSL), no prior work has explored their use for pose estimation. These methods rely on low rank-based adaptation approaches, which constrain model updates to a low-rank space. We propose Endo-FASt3r, the first monocular SSL depth and pose estimation framework that uses foundation models for both tasks. We extend the Reloc3r relative pose estimation foundation model by designing Reloc3rX, introducing modifications necessary for convergence in SSL. We also present DoMoRA, a novel adaptation technique that enables higher-rank updates and faster convergence. Experiments on the SCARED dataset show that Endo-FASt3r achieves a substantial $10\%$ improvement in pose estimation and a $2\%$ improvement in depth estimation over prior work. Similar performance gains on the Hamlyn and StereoMIS datasets reinforce the generalisability of Endo-FASt3r across different datasets.

内窥镜自监督学习位姿估计3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。