arXiv:2501.18124cs.CVcs.AI2025-01ICRA

通过多模态特征学习实现多种内窥镜的实时运动追踪

REMOTE: Real-time Ego-motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning

  • 融合光流、场景特征与双帧联合特征进行相对位姿预测
  • 在三个数据集上精度超越现有方法,推理速度超30帧/秒
  • 适合需要实时导航的手术机器人与内窥镜自动化系统

内窥镜的实时自运动追踪对于高效导航和机器人自动化至关重要。本文提出一种新框架,实现多种内窥镜的实时自运动追踪。首先,设计一个多模态视觉特征学习网络,利用光流运动特征、场景特征以及两帧相邻观测的联合特征进行相对位姿预测。由于拼接图像在通道维度具有更强的相关性,提出一种基于注意力机制的新特征提取器,以融合多维信息。为进一步提取完整的特征表示,设计了一种新型姿态解码器,从框架末端的拼接特征图中预测姿态变换。最后,基于相对位姿计算内窥镜的绝对位姿。在三个不同内窥镜场景的数据集上进行实验,结果表明所提方法优于现有最先进方法。此外,该方法推理速度超过30帧每秒,满足实时性要求。

原文摘要 · Abstract (English)

Real-time ego-motion tracking for endoscope is a significant task for efficient navigation and robotic automation of endoscopy. In this paper, a novel framework is proposed to perform real-time ego-motion tracking for endoscope. Firstly, a multi-modal visual feature learning network is proposed to perform relative pose prediction, in which the motion feature from the optical flow, the scene features and the joint feature from two adjacent observations are all extracted for prediction. Due to more correlation information in the channel dimension of the concatenated image, a novel feature extractor is designed based on an attention mechanism to integrate multi-dimensional information from the concatenation of two continuous frames. To extract more complete feature representation from the fused features, a novel pose decoder is proposed to predict the pose transformation from the concatenated feature map at the end of the framework. At last, the absolute pose of endoscope is calculated based on relative poses. The experiment is conducted on three datasets of various endoscopic scenes and the results demonstrate that the proposed method outperforms state-of-the-art methods. Besides, the inference speed of the proposed method is over 30 frames per second, which meets the real-time requirement. The project page is here: remote-bmxs.netlify.app

内窥镜姿态追踪实时系统多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。