arXiv:2506.16017cs.CVcs.RO2025-06中稿 · IROS 2025被引 1

端镜视觉中,分步自监督训练提升单目深度估计精度。

EndoMUST: Monocular Depth Estimation for Robotic Endoscopy via End-to-end Multi-step Self-supervised Training

  • 分三步训练:光流对齐、多尺度分解、变换对齐,每步专注特定任务。
  • 在SCARED数据集上误差降低4%~10%,零样本迁移至Hamlyn表现领先。
  • 适合做内窥镜机器人导航与深度感知的开发者参考。

单目深度估计与自身运动估计在稳定、准确、高效的机器人辅助内窥镜中具有重要意义。为应对内窥镜场景中的光照变化与纹理稀疏问题,现有方法引入了光流、外观流及内在图像分解等技术。然而,多模块的有效训练策略仍是自监督深度估计面临的挑战。本文提出一种端到端多步自监督训练框架,每轮训练分为三个步骤:光流配准、多尺度图像分解和多变换对齐,每步仅训练相关网络,避免无关信息干扰。基于基础模型的参数高效微调,该方法在SCARED数据集上实现自监督深度估计的当前最优性能,在Hamlyn数据集上实现零样本深度估计,误差降低4%~10%。代码已开源:https://github.com/BaymaxShao/EndoMUST。

原文摘要 · Abstract (English)

Monocular depth estimation and ego-motion estimation are significant tasks for scene perception and navigation in stable, accurate and efficient robot-assisted endoscopy. To tackle lighting variations and sparse textures in endoscopic scenes, multiple techniques including optical flow, appearance flow and intrinsic image decomposition have been introduced into the existing methods. However, the effective training strategy for multiple modules are still critical to deal with both illumination issues and information interference for self-supervised depth estimation in endoscopy. Therefore, a novel framework with multistep efficient finetuning is proposed in this work. In each epoch of end-to-end training, the process is divided into three steps, including optical flow registration, multiscale image decomposition and multiple transformation alignments. At each step, only the related networks are trained without interference of irrelevant information. Based on parameter-efficient finetuning on the foundation model, the proposed method achieves state-of-the-art performance on self-supervised depth estimation on SCARED dataset and zero-shot depth estimation on Hamlyn dataset, with 4\%$\sim$10\% lower error. The evaluation code of this work has been published on https://github.com/BaymaxShao/EndoMUST.

深度估计内窥镜自监督机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。