arXiv:2603.08324cs.ROcs.AI2026-03

提出视觉导航新方法,解决内窥机器人在复杂腔道中定位不准问题。

EndoSERV: A Vision-based Endoluminal Robot Navigation System

  • 分段+虚实映射,提升复杂腔道下的定位精度
  • 无需真实姿态标签,仍可实现高精度导航
  • 适合临床内窥手术机器人,尤其在缺乏标记物场景

机器人辅助的内腔操作正广泛用于早期癌症干预,但腔道结构狭窄、曲折且易受组织变形和体内伪影影响,导致导航困难。现有视觉定位方法因缺乏显著特征而易出错。本文提出一种新型端到端视觉导航方法EndoSERV,包含两个核心模块: - SE(Segment-to-structure):将长程复杂腔道分割为子段,独立估计位姿; - RV(Real-to-Virtual):通过高效迁移学习,将真实图像特征映射至虚拟域,利用虚拟姿态真值进行训练。 EndoSERV采用离线预训练提取与纹理无关的特征,并在在线阶段适应真实环境。基于公开数据集和临床数据集的大量实验表明,该方法即使未使用任何真实姿态标签,也能实现高精度导航,显著优于现有方法。

原文摘要 · Abstract (English)

Robot-assisted endoluminal procedures are increasingly used for early cancer intervention. However, the intricate, narrow and tortuous pathways within the luminal anatomy pose substantial difficulties for robot navigation. Vision-based navigation offers a promising solution, but existing localization approaches are error-prone due to tissue deformation, in vivo artifacts and a lack of distinctive landmarks for consistent localization. This paper presents a novel EndoSERV localization method to address these challenges. It includes two main parts, \textit{i.e.}, \textbf{SE}gment-to-structure and \textbf{R}eal-to-\textbf{V}irtual mapping, and hence the name. For long-range and complex luminal structures, we divide them into smaller sub-segments and estimate the odometry independently. To cater for label insufficiency, an efficient transfer technique maps real image features to the virtual domain to use virtual pose ground truth. The training phases of EndoSERV include an offline pretraining to extract texture-agnostic features, and an online phase that adapts to real-world conditions. Extensive experiments based on both public and clinical datasets have been performed to demonstrate the effectiveness of the method even without any real pose labels.

内窥机器人视觉导航医学影像虚实映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。