用立体可微渲染实现手术机器人无标记实时追踪,精度媲美有标记方案。
Streamlining stereo differentiable rendering for marker-free real-time tracking of surgical robots

- 通过帧间优化与动态超参数调节,实现连续姿态估计。
- 1080p下30帧/秒实时追踪,误差仅1.7厘米/0.6度,遮挡下仍达1.2厘米均值误差。
- 适合需要高精度、无标记、多机器人协同的手术导航系统开发。
目的:手术室中基于标记的机器人追踪易受遮挡影响。本文评估立体可微渲染在无标记、实时机器人位姿追踪中的应用,有望提升安全性、缩短准备时间并支持多机器人协作。方法:将无标记位姿估计框架roboreg扩展为在线动态追踪,采用(i)帧间传播的姿态估计与运动自适应超参数调优,及(ii)CUDA流并行化分割与优化,并结合CUDA图加速分割。在38段无遮挡和5段遮挡位移序列上进行评估,使用静态起止真值标定和动态标记基参考追踪。结果:实现1080p下30帧/秒实时追踪(较原版roboreg的14帧/秒显著提升),精度达1.7厘米/0.6度(对比静态真值),27,460帧平均3D误差1.2厘米(1,242帧遮挡下为1.53厘米)。相较FoundationPose,在动态估计上提升11%(遮挡下达63%),静态估计提升250%,推理速度提升6倍。结论:立体可微渲染实现了高分辨率、实时、无标记的手术机器人追踪,性能媲美标记法且优于基础模型基准。
原文摘要 · Abstract (English)
Purpose: Marker-based tracking of surgical robots is occlusion-prone in cluttered operating rooms. We evaluate stereo differentiable rendering for marker-free, real-time robot pose tracking, potentially improving safety, reducing setup time, and enabling multi-robot interaction. Methods: We extend the markerless pose estimation framework roboreg to online dynamic tracking via (i) sequential optimisation that propagates pose estimates across frames with motion-adaptive hyperparameter tuning, and (ii) CUDA stream parallelisation of segmentation and optimisation, combined with CUDA-graph accelerated segmentation. We evaluate on 38 unobstructed and 5 occluded displacement sequences with static start/end ground-truth calibrations and dynamic marker-based reference tracking. Results: We achieve real-time 1080p tracking at 30 fps (up from 14 fps for vanilla roboreg), matching the camera frame rate. Accuracy reaches 1.7 cm / 0.6 deg against static ground truth and 1.2 cm mean 3D error over 27,460 frames against the marker-based reference (1.53 cm over 1,242 occluded frames). Our method outperforms FoundationPose by 11% in dynamic estimation (63% under occlusion) and 250% in static estimation, with 6x faster inference. Conclusions: Stereo differentiable rendering enables real-time, high-resolution marker-free surgical robot tracking, on par with marker-based approaches and surpassing foundation-model baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。