arXiv:2508.11200cs.ROcs.AI2025-08被引 2

用世界模型实现无需训练的手术器械抓取,适配多种新物体和干扰。

Visuomotor Grasping with World Models for Surgical Robots

  • 基于世界模型与单双目视觉,学习通用抓取策略。
  • 真实场景下抓取成功率65%,支持未见物体与夹具。
  • 适合追求高鲁棒性、免重训的外科机器人研究者。

抓取是机器人辅助手术(RAS)中的基础任务,自动化可减轻医生负担并提升效率、安全性和一致性。现有方法依赖显式物体位姿追踪或手工视觉特征,难以泛化到新物体、抗视觉干扰,且难处理变形物体。本文针对三个挑战:(i) 将视觉-运动策略从仿真迁移至离体手术场景;(ii) 仅使用标准双目内窥镜进行视觉感知;(iii) 采用单一无对象特性的策略,泛化于多样未见手术物体。提出手术抓取通用框架GASv2,结合世界模型架构、手术感知流水线及混合控制机制。通过领域随机化在仿真中训练,并部署于真实机器人,在模拟与离体实验中均实现65%抓取成功率,成功泛化至未知物体与夹具,适应多种扰动,验证了其性能、泛化性与鲁棒性。

原文摘要 · Abstract (English)

Grasping is a fundamental task in robot-assisted surgery (RAS), and automating it can reduce surgeon workload while enhancing efficiency, safety, and consistency beyond teleoperated systems. Most prior approaches rely on explicit object pose tracking or handcrafted visual features, limiting their generalization to novel objects, robustness to visual disturbances, and the ability to handle deformable objects. Visuomotor learning offers a promising alternative, but deploying it in RAS presents unique challenges, such as low signal-to-noise ratio in visual observations, demands for high safety and millimeter-level precision, as well as the complex surgical environment. This paper addresses three key challenges: (i) sim-to-real transfer of visuomotor policies to ex vivo surgical scenes, (ii) visuomotor learning using only a single stereo camera pair -- the standard RAS setup, and (iii) object-agnostic grasping with a single policy that generalizes to diverse, unseen surgical objects without retraining or task-specific models. We introduce Grasp Anything for Surgery V2 (GASv2), a visuomotor learning framework for surgical grasping. GASv2 leverages a world-model-based architecture and a surgical perception pipeline for visual observations, combined with a hybrid control system for safe execution. We train the policy in simulation using domain randomization for sim-to-real transfer and deploy it on a real robot in both phantom-based and ex vivo surgical settings, using only a single pair of endoscopic cameras. Extensive experiments show our policy achieves a 65% success rate in both settings, generalizes to unseen objects and grippers, and adapts to diverse disturbances, demonstrating strong performance, generality, and robustness.

手术机器人视觉-运动世界模型泛化抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。