arXiv:2512.18068cs.RO2025-12被引 4

从单目手术视频中估算器械运动参数,让机器人学习手术动作不再依赖真实数据。

SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning

  • 基于可微渲染优化器械位姿,从单张视频还原运动轨迹与关节角度。
  • 用估计的运动数据训练的机器人策略成功率接近真实数据训练结果。
  • 适合想用公开手术视频训练机器人手术模型的研究者。

模仿学习在实现自主灵巧操作方面展现出巨大潜力,包括学习外科手术任务。要充分发挥其在手术中的应用潜力,需要临床数据集支持,但现有数据通常缺乏当前模仿学习方法所需的运动学数据。一个有前景的大规模手术示范来源是网络上可获取的单目手术视频,因此从单目视频中进行姿态估计成为实现大规模机器人学习的关键步骤。为此,我们提出SurgiPose,一种基于可微渲染的方法,用于从单目手术视频中估计运动学信息,无需直接访问真实运动学数据。该方法通过优化器械位姿参数,使渲染图像与真实图像之间的差异最小化,从而推断工具轨迹和关节角度。我们在两个机器人手术任务(组织提起和针拾取)上使用da Vinci Research Kit Si(dVRK Si)进行实验,分别用真实测量的运动学数据和视频估计的运动学数据训练模仿学习策略,并比较其性能。结果显示,基于估计数据训练的策略取得了与真实数据训练相当的成功率,证明了基于单目视频的运动学估计在手术机器人学习中的可行性。本工作为利用在线手术数据大规模学习自主手术策略奠定了基础。

原文摘要 · Abstract (English)

Imitation learning (IL) has shown immense promise in enabling autonomous dexterous manipulation, including learning surgical tasks. To fully unlock the potential of IL for surgery, access to clinical datasets is needed, which unfortunately lack the kinematic data required for current IL approaches. A promising source of large-scale surgical demonstrations is monocular surgical videos available online, making monocular pose estimation a crucial step toward enabling large-scale robot learning. Toward this end, we propose SurgiPose, a differentiable rendering based approach to estimate kinematic information from monocular surgical videos, eliminating the need for direct access to ground truth kinematics. Our method infers tool trajectories and joint angles by optimizing tool pose parameters to minimize the discrepancy between rendered and real images. To evaluate the effectiveness of our approach, we conduct experiments on two robotic surgical tasks: tissue lifting and needle pickup, using the da Vinci Research Kit Si (dVRK Si). We train imitation learning policies with both ground truth measured kinematics and estimated kinematics from video and compare their performance. Our results show that policies trained on estimated kinematics achieve comparable success rates to those trained on ground truth data, demonstrating the feasibility of using monocular video based kinematic estimation for surgical robot learning. By enabling kinematic estimation from monocular surgical videos, our work lays the foundation for large scale learning of autonomous surgical policies from online surgical data.

手术机器人模仿学习单目估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。