arXiv:2604.21017cs.ROcs.AI2026-04被引 8

构建首个大规模开源医疗机器人数据集,推动医学机器人基础模型发展。

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics

论文配图:Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics
图 1 · 摘自论文原文
  • 收集50+机构多平台数据,含视频与运动学同步信息
  • 实现25%任务完成率,64%缝合成功率,创基准纪录
  • 支持跨机器人仿真与政策评估,适合医疗机器人研究者

自主医疗机器人有望提升患者疗效、减轻医护人员负担、普及医疗服务并实现超人精度。然而,当前发展受限于根本性数据瓶颈:现有医疗机器人数据集规模小、单一设备、极少开放共享,制约了基础模型的构建。本文推出Open-H-Embodiment,迄今最大且公开的医疗机器人视频与运动学同步数据集,覆盖超过50家机构及多种机器人平台,包括CMR Versius、Intuitive Surgical da Vinci、da Vinci Research Kit(dVRK)、Rob Surgical BiTrack、Virtual Incision MIRA、Moon Surgical Maestro及多种定制系统,涵盖手术操作、机器人超声与内窥镜等任务。我们基于该数据集展示了两项基础模型成果:GR00T-H是首个开源的医疗机器人视觉-语言-动作基础模型,在结构化缝合基准上唯一实现端到端任务完成(25%成功率,其余为0%),并在29步离体缝合序列中达到64%平均成功率。此外,我们训练了Cosmos-H-Surgical-Simulator,首个可实现多平台手术模拟的动作条件世界模型,仅需单个检查点即可覆盖九种机器人平台,支持虚拟策略评估与领域内合成数据生成。结果表明,开放的大规模医疗机器人数据采集可成为研究社区的关键基础设施,推动机器人学习与世界建模等方向进步。

原文摘要 · Abstract (English)

Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem: existing medical robotic datasets are small, single-embodiment, and rarely shared openly, restricting the development of foundation models that the field needs to advance. We introduce Open-H-Embodiment, the largest open dataset of medical robotic video with synchronized kinematics to date, spanning more than 50 institutions and multiple robotic platforms including the CMR Versius, Intuitive Surgical's da Vinci, da Vinci Research Kit (dVRK), Rob Surgical BiTrack, Virtual Incision's MIRA, Moon Surgical Maestro, and a variety of custom systems, spanning surgical manipulation, robotic ultrasound, and endoscopy procedures. We demonstrate the research enabled by this dataset through two foundation models. GR00T-H is the first open foundation vision-language-action model for medical robotics, which is the only evaluated model to achieve full end-to-end task completion on a structured suturing benchmark (25% of trials vs. 0% for all others) and achieves 64% average success across a 29-step ex vivo suturing sequence. We also train Cosmos-H-Surgical-Simulator, the first action-conditioned world model to enable multi-embodiment surgical simulation from a single checkpoint, spanning nine robotic platforms and supporting in silico policy evaluation and synthetic data generation for the medical domain. These results suggest that open, large-scale medical robot data collection can serve as critical infrastructure for the research community, enabling advances in robot learning, world modeling, and beyond.

医疗机器人基础模型数据集世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。