首次评估模仿学习在开腹手术缝合中的应用,$π_0$模型表现最优。
Imitation Learning for Robot Assistance in Open Surgery: A Multi-Policy Evaluation on Suture Following

- 采用四种不同架构的模仿学习策略,基于32,374帧数据训练
- 理想条件下任务成功率50%-75%,深度误差是主要失败原因
- $π_0$模型数据效率高、对背景变化鲁棒,适合临床手术流程
本研究首次评估了通用模仿学习在开腹手术中人机协作辅助的可行性,聚焦缝合过程中的抓-拉-放动作。在开源机械臂上收集160次远程操作演示(共32,374帧),在28个模型、32种配置下,基于数据集规模、相机视角和背景变化三个临床相关维度,对比了四种架构差异显著的模仿学习策略(ACT、Diffusion Policy、SmolVLA、$π_0$)。结果表明,在理想条件下,四类策略任务成功率可达50%-75%,深度误差是各类架构的主要失败模式。其中,$π_0$凭借预训练视觉-语言骨干网络,展现出更强的数据效率、对背景变化的鲁棒性以及与手术流程兼容的平滑轨迹。在真实人机缝合实验中,$π_0$实现92%的缝合完成率。研究证实开腹手术协作机器人辅助具备模仿学习可行性,并指出深度感知与末端执行器设计是临床转化的关键优先事项。
原文摘要 · Abstract (English)
This study presents the first evaluation of general-purpose imitation learning for surgeon-robot collaborative assistance in open surgery, targeting suture following: the grab-pull-release motion an assistant performs at every stitch. We collect 160 teleoperated demonstrations (32,374 frames) on an open-source robot arm, benchmark four architecturally diverse imitation learning policies (ACT, Diffusion Policy, SmolVLA, $π_0$) across 28 trained models evaluated in 32 configurations along three clinically motivated dimensions: dataset size, camera viewpoint, and background variation. Our results demonstrate that under ideal conditions, the four policies achieve $50$-$75\%$ task success, with depth error as the dominant failure mode across all architectures. Among all policies, $π_0$ achieves the strongest results with a pretrained vision-language backbone, demonstrating superior data efficiency, greater robustness to background variation, and smoother trajectories compatible with surgical workflow. When deployed in a surgeon-robot suturing trial, $π_0$ yields a $92\%$ stitch completion rate. These findings establish collaborative robotic assistance in open surgery as a feasible target for imitation learning and highlight depth perception and end-effector design as key priorities for clinical translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。