arXiv:2601.21971cs.ROcs.AI2026-01中稿 · Robotics:Science a…被引 5

用轻量级模型仅靠150次演示,实现手术抓取与牵拉的高精度控制

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

  • 采用监督式混合专家架构,提升视觉-动作映射能力
  • 在少于150次示范下达成92%成功率,且抗干扰能力强
  • 可零样本迁移到活体组织,适合临床前机器人手术部署

模仿学习在机器人操作中已取得显著成果,但其在手术机器人的应用仍受限于数据稀缺、空间受限及对安全性和可预测性的极高要求。本文提出一种面向阶段化手术操作任务的监督式混合专家(MoE)架构,可嵌入任意自主策略之上。与依赖多摄像头或数千次示范的现有方法不同,我们证明:当结合该架构时,轻量级动作解码器如动作分块变换器(ACT),仅需少于150次示范,即可利用双目内窥镜图像学会复杂、长时程的操作。我们在肠管抓取与牵拉这一协作手术任务上评估该方法,机器人助手需根据主刀医生的视觉提示,对可变形组织进行精准抓取并持续牵拉。结果显示,通用视觉语言动作模型即使在分布内条件下也未能完整掌握任务;尽管标准ACT在分布内表现中等,但引入监督式MoE后,其性能显著提升,不仅在分布内成功率更高,且在分布外场景(如新抓取位置、光照降低、部分遮挡)中表现出更强鲁棒性。值得注意的是,该方法能泛化至未见过的观测视角,并在无需额外训练的情况下零样本迁移至离体猪组织,为活体应用提供可行路径。为此,我们还展示了在活体猪手术中策略执行的定性初步结果。

原文摘要 · Abstract (English)

Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due to data scarcity, constrained workspaces, and the need for an exceptional level of safety and predictability. We present a supervised Mixture-of-Experts (MoE) architecture designed for phase-structured surgical manipulation tasks, which can be added on top of any autonomous policy. Unlike prior surgical robot learning approaches that rely on multi-camera setups or thousands of demonstrations, we show that a lightweight action decoder policy like Action Chunking Transformer (ACT) can learn complex, long-horizon manipulation from less than 150 demonstrations using solely stereo endoscopic images, when equipped with our architecture. We evaluate our approach on the collaborative surgical task of bowel grasping and retraction, where a robot assistant interprets visual cues from a human surgeon, executes targeted grasping on deformable tissue, and performs sustained retraction. Our results show that generalist Vision Language Action models fail to acquire the task entirely, even under standard in-distribution conditions. Furthermore, while standard ACT achieves moderate success in-distribution, adopting a supervised MoE architecture significantly boosts its performance, yielding higher success rates in-distribution and demonstrating superior robustness in out-of-distribution scenarios, including novel grasp locations, reduced illumination, and partial occlusions. Notably, it generalizes to unseen testing viewpoints and also transfers zero-shot to ex vivo porcine tissue without additional training, offering a promising pathway toward in vivo deployment. To support this statement, we present qualitative preliminary results of policy roll-outs during in vivo porcine surgery.

手术机器人模仿学习混合专家零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。