用开源手术视频提升机器人仿习学习泛化能力,解决真实组织数据稀缺问题。
SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos

- 融合带动作标签的模拟数据与开源手术视频,用近似运动信息作弱监督
- 在达芬奇机器人上实现针头抓取和胆囊切除切割任务,显著提升真实组织泛化性
- 适合追求真实场景通用性的医疗机器人研究者,可扩展至多种手术任务
基于学习的外科机器人自主系统需要大规模同步视频与机器人动作演示,但临床或真实组织环境下此类数据极为稀少,因机器人运动学数据通常无法在非受控研究系统中获取。而研究平台上的模拟数据虽有精确动作标签,却缺乏真实组织的视觉多样性。本文提出SurgVIL框架,利用开源手术视频扩展外科机器人模仿学习。该框架结合带有运动标注的模拟机器人演示与来自开源数据集及网络来源的手术视频进行策略学习。由于这些视频无机器人运动标签,我们通过估计近似运动作为弱监督信号。在达芬奇机器人两个任务(针头抓取与胆囊切除切割)上评估,使用ACT、π₀与GR00T-H三种模型架构均表明,加入手术视频显著提升模型在真实组织和分布外场景的泛化能力,验证了从模拟训练迈向通用外科机器人策略的可扩展路径。
原文摘要 · Abstract (English)
Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot actions, but such data are exceedingly rare in clinical or realistic tissue settings because robot kinematics are typically inaccessible outside controlled research systems. In contrast, phantom data collected on research platforms provide accurate action labels but lack the visual diversity of real tissue. We propose SurgVIL, a framework for scaling surgical robot imitation learning using open-source surgical videos. SurgVIL combines kinematically labeled phantom robot demonstrations with surgical videos from open-source datasets and online sources for policy learning. Since these videos lack robot motion labels, we estimate approximate kinematics as weak supervision. We evaluate SurgVIL on two da Vinci robot tasks: needle pick-up and cholecystectomy cutting. Across ACT, $π_0$, and GR00T-H backbones, adding surgical videos substantially improves generalization to real-tissue and out-of-distribution settings, suggesting a scalable path from phantom training toward generalizable surgical robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。