用可微分仿真训练视觉模型,实时纠正手术机器人工具位姿误差
Real-time Capable Learning-based Visual Tool Pose Correction via Differentiable Simulation
- 基于视觉变换器,通过可微分物理仿真端到端训练
- 误差降低超50%,达到优化方法性能,推理速度达22赫兹
- 无需标注即可零样本泛化,微调后性能更优
机器人辅助微创手术的自主化有望减轻外科医生的认知与操作负担,提升手术效率。然而,由于末端执行器本体感知精度差,实现精准自主控制仍具挑战。关节编码器读数常因缆绳传动中的运动学非理想性而失准。基于视觉的姿态估计虽有效,但缺乏实时性、泛化能力或难以训练。本文提出一种基于视觉变换器的实时姿态估计方法,利用端到端可微分运动学与渲染进行训练。在真实机器人数据集上验证了其修正噪声姿态估计的能力及近实时处理潜力。该方法将手眼转换误差降低超过50%,性能媲美现有优化方法;推理速度提升四倍,可达22赫兹。在未见数据集上实现零样本预测,展现良好泛化能力,且可通过无标签微调进一步提升性能。
原文摘要 · Abstract (English)
Autonomy in robot-assisted minimally invasive surgery has the potential to reduce surgeon cognitive and task load, thereby increasing procedural efficiency. However, implementing accurate autonomous control can be difficult due to poor end-effector proprioception. Joint encoder readings are typically inaccurate due to kinematic non-idealities in their cable-driven transmissions. Vision-based pose estimation approaches are highly effective, but lack real-time capability, generalizability, or can be hard to train. In this work, we demonstrate a real-time capable, Vision Transformer-based pose estimation approach that is trained using end-to-end differentiable kinematics and rendering. We demonstrate the potential of this approach to correct for noisy pose estimates through a real robot dataset and the potential real-time processing ability. Our approach is able to reduce more than 50% of hand-eye translation errors in the dataset, reaching the same performance level as an existing optimization-based method. Our approach is four times faster, and capable of near real-time inference at 22 Hz. A zero-shot prediction on an unseen dataset shows good generalization ability, and can be further finetuned for increased performance without human labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。