arXiv:2510.03532cs.ROcs.CV2025-10被引 4

统一检测关键点与边缘,实现手术机器人实时精准位姿估计

Efficient Surgical Robotic Instrument Pose Reconstruction in Real World Conditions Using Unified Feature Detection

  • 共享编码架构同步检测关键点与杆状边缘
  • 单次推理完成检测,推理速度满足在线控制需求
  • 基于合成数据训练,在复杂手术场景达顶尖精度

视觉引导的机器人控制系统中,相机到机器人的精确标定至关重要,尤其在微创手术机器人中,仪器需进行高精度微操作。然而,微创手术机器人具有长运动链且部分自由度在摄像头中可见性差,传统假设刚性结构且可视性良好的标定方法难以适用。此前工作尝试基于特征点或渲染的方法应对真实场景挑战,但普遍存在特征检测不一致或推理时间过长的问题,不适合在线控制。本文提出一种新框架,通过共享编码统一检测几何基元(关键点与杆状边缘),利用投影几何实现高效位姿估计。该架构在一次推理中同时检测关键点与边缘,并在大规模合成数据上以投影标签进行训练。实验在特征检测与位姿估计两个层面评估,定性和定量结果均表明该方法在复杂手术环境中具备快速性能和领先精度。

原文摘要 · Abstract (English)

Accurate camera-to-robot calibration is essential for any vision-based robotic control system and especially critical in minimally invasive surgical robots, where instruments conduct precise micro-manipulations. However, MIS robots have long kinematic chains and partial visibility of their degrees of freedom in the camera, which introduces challenges for conventional camera-to-robot calibration methods that assume stiff robots with good visibility. Previous works have investigated both keypoint-based and rendering-based approaches to address this challenge in real-world conditions; however, they often struggle with consistent feature detection or have long inference times, neither of which are ideal for online robot control. In this work, we propose a novel framework that unifies the detection of geometric primitives (keypoints and shaft edges) through a shared encoding, enabling efficient pose estimation via projection geometry. This architecture detects both keypoints and edges in a single inference and is trained on large-scale synthetic data with projective labeling. This method is evaluated across both feature detection and pose estimation, with qualitative and quantitative results demonstrating fast performance and state-of-the-art accuracy in challenging surgical environments.

机器人控制姿态估计医学影像深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。