用视频预训练提升手术机器人数据效率,仅靠少量标注动作就能更好完成复杂操作。
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

- 构建视觉-动作联合模型,从无标注视频中学习手术场景动态
- 在4个仿真任务上成功率从63.5%提升至77.8%,接触密集任务提升20个百分点
- 适合资源有限、需少样本高效训练的手术机器人研究者
手术机器人操控策略的学习受限于带动作标签示范数据的稀缺:同步视频-运动学轨迹(如dVRK)收集成本高昂,而手术任务需要精确接触处理、长时序推理和双臂协同。内窥镜视频相比而言更易获取且丰富,自然可用来学习手术场景的世界模型。但现有模型多将视频用于仿真或策略评估,很少将其动态知识转化为闭环控制。本研究提出外科世界-动作模型(Surgical WAM),基于Cosmos Policy构建统一生成模型,联合预测未来内窥镜观测与可执行的机器人动作片段。Surgical WAM先从无动作标注视频中学习视觉动态,再在固定数量的动作标注数据上微调;部署时作为闭环滚动优化控制器,执行预测动作片段的前缀并根据新观测重新规划。在四个仿真手术任务上,视频预训练使平均成功率从63.5%提升至77.8%,其中佩吉转移任务绝对提升20个百分点,接触密集和双臂任务改善最显著。结果表明,无动作视频提供了可迁移的视觉动态先验,使在有限动作监督下学习手术机器人控制成为可能,验证了数据高效视频预训练是规模化手术机器人学习的有效路径。
原文摘要 · Abstract (English)
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to synchronized video--kinematics trajectories, and a natural way to exploit it is to learn world models of surgical scenes. However, existing surgical world models use video primarily for simulation or policy evaluation, and rarely translate the learned dynamics into closed-loop control. This gap raises our central question: under a fixed budget of action-labeled demonstrations, does action-free video pretraining improve closed-loop surgical manipulation? To answer it, we introduce the Surgical World-Action Model (Surgical WAM), a unified generative model built on Cosmos Policy that jointly predicts future endoscopic observations and executable surgical robot action chunks. Surgical WAM first learns surgical visual dynamics from action-free video and is then fine-tuned on the fixed action-labeled budget; at deployment, it acts as a closed-loop, receding-horizon controller that executes a short prefix of each predicted action chunk and replans from the resulting observation. On a suite of four simulated surgical manipulation tasks, video pretraining improves the average success rate from 63.5% to 77.8%, including an absolute gain of 20 percentage points on PegTransfer, with the largest improvements on contact-rich and bimanual tasks. These results demonstrate that action-free video provides transferable visual dynamics priors for learning surgical robot control with limited action supervision, positioning data-efficient video pretraining as a practical path toward scaling up surgical robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。