用合成视频训练手术机器人,突破数据稀缺瓶颈
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
- 构建世界模型生成逼真手术视频,填补数据空白
- 通过逆动力学模型推断伪运动轨迹,实现视频-动作配对
- 在真实机器人上验证,性能超越仅依赖真实数据的模型
数据稀缺是实现全自动手术机器人的主要障碍。尽管大规模视觉-语言-动作(VLA)模型在家庭和工业操作中展现出强大泛化能力,得益于跨领域的视频-动作配对数据,但手术机器人缺乏包含视觉观测与精确机器人运动学的数据集。相比之下,大量手术视频存在,却缺少对应的动作标签,无法直接用于模仿学习或VLA训练。本文提出Cosmos-H-Surgical,一个专为手术物理人工智能设计的世界模型。我们构建了手术动作文本对齐(SATA)数据集,提供针对手术机器人的详细动作描述。基于最先进物理AI世界模型与SATA,我们训练出能生成多样化、可泛化且逼真的手术视频的模型。首次利用逆动力学模型从合成手术视频中推断伪运动学数据,生成合成的视频-动作配对数据。实验表明,使用这些增强数据训练的手术VLA策略,在真实手术机器人平台上显著优于仅使用真实演示训练的模型。该方法通过利用海量无标签手术视频和生成式世界建模,为自主手术技能获取提供了可扩展路径,推动通用且数据高效的手术机器人策略发展。
原文摘要 · Abstract (English)
Data scarcity remains a fundamental barrier to achieving fully autonomous surgical robots. While large scale vision language action (VLA) models have shown impressive generalization in household and industrial manipulation by leveraging paired video action data from diverse domains, surgical robotics suffers from the paucity of datasets that include both visual observations and accurate robot kinematics. In contrast, vast corpora of surgical videos exist, but they lack corresponding action labels, preventing direct application of imitation learning or VLA training. In this work, we aim to alleviate this problem by learning policy models from Cosmos-H-Surgical, a world model designed for surgical physical AI. We curated the Surgical Action Text Alignment (SATA) dataset with detailed action description specifically for surgical robots. Then we built Cosmos-H-Surgical based on the most advanced physical AI world model and SATA. It's able to generate diverse, generalizable and realistic surgery videos. We are also the first to use an inverse dynamics model to infer pseudokinematics from synthetic surgical videos, producing synthetic paired video action data. We demonstrate that a surgical VLA policy trained with these augmented data significantly outperforms models trained only on real demonstrations on a real surgical robot platform. Our approach offers a scalable path toward autonomous surgical skill acquisition by leveraging the abundance of unlabeled surgical video and generative world modeling, thus opening the door to generalizable and data efficient surgical robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。