无需标记物,用双目视觉实现手术机器人末端6自由度精准定位。
MarUco: A Markerless 6D Pose Estimation Framework for Closed-Loop Control of Surgical Continuum Manipulators
- 基于仿真生成大量带标注数据,结合多特征融合网络初估姿态。
- 单次渲染比对模块精修姿态,52.4毫秒完成一帧处理,误差仅0.78毫米。
- 无需标定,自适应真实场景,适合手术机器人闭环控制应用。
柔性内窥镜连续体机械臂具有高灵活性和复杂解剖结构的可达性,但非线性迟滞限制了前馈控制精度。闭环控制可补偿此类误差,但需精确的六自由度(6D)末端位姿反馈。本文提出MarUco,一种仅使用操作期间双目视觉的无标记6D位姿估计框架。通过光栅拟刚体仿真流程生成大规模无手动标注的训练数据。多特征融合网络整合双视角的掩码、关键点、热力图和边界框信息以估计初始位姿,随后采用学习型单次渲染-比对模块在不进行迭代优化的情况下精修位姿。基于运动学的相机到机器人基座外参估计方法,以及利用未标注真实双目图像的自监督适配策略,有效缓解了仿真到现实的位姿误差。在1000个真实样本上,MarUco实现0.78±0.50毫米的平移误差和3.07±1.25°的旋转误差,每对双目图像处理总时长为52.4毫秒。在八个参考路径的闭环实验中,终端平移误差平均为1.8毫米,相较未补偿的开环控制降低88%。据我们所知,这是首个面向连续体机械臂的位置式视觉伺服无标记框架。
原文摘要 · Abstract (English)
Flexible endoscopic continuum manipulators offer high dexterity and access to complex anatomy, but nonlinear hysteresis limits feedforward control accuracy. Closed-loop control can compensate for these errors but requires accurate six-degree-of-freedom (6D) end-effector pose feedback. We present MarUco, a markerless 6D pose estimation framework for closed-loop control using only stereo vision during operation. A photorealistic pseudo-rigid-body simulation pipeline generates large-scale annotated training data without manual labeling. A multifeature fusion network integrates masks, keypoints, heatmaps, and bounding boxes from both stereo views to estimate an initial pose, followed by a learned single-pass render-and-compare module that refines the pose without iterative optimization. Kinematics-free hand-eye calibration estimates camera-to-robot-base extrinsics, and self-supervised adaptation uses unlabeled real stereo pairs to mitigate sim-to-real pose error. Across 1,000 real samples, MarUco achieves translation and rotation errors of 0.78 $\pm$ 0.50 mm and 3.07 $\pm$ 1.25°, respectively, with a total processing time of 52.4 ms per stereo pair. In closed-loop experiments over eight reference paths, MarUco achieves a mean terminal translation error of 1.8 mm, an 88% reduction relative to uncompensated open-loop control. To the best of our knowledge, this is the first markerless position-based visual servoing framework for continuum manipulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。