arXiv:2505.11920cs.RO2025-05被引 24

将人类视频转为机器人视角数据,提升机器人预训练效果。

H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos

  • 用人体手势估计与仿真机械臂动作重定向,生成机器人视角视频。
  • 在仿真和真实场景中均提升任务成功率1.3%至23.3%。
  • 适合做机器人视觉-语言-动作模型预训练,尤其关注跨机体泛化。

利用第一人称人类视频进行大规模机器人预训练已证明有效,但因人类手部与机器人之间的显著视觉差异,导致预训练模型性能受限。为此,我们提出H2R:一种从人类视频生成机器人视角数据的增强流程。H2R从视频中估计人体手部姿态,将其动作重定向至仿真机械臂,通过分割与修复移除人类肢体,并将渲染的机器人模型以相机对齐几何方式融合回原帧。该过程在预训练阶段显式弥合了人机形态间的视觉差距。我们将H2R应用于Ego4D和SSv2等大规模第一人称视频数据集。为验证增强效果,引入基于CLIP的图文相似性度量,定量评估渲染帧与原始人类动作的语义一致性。在仿真与真实环境中全面评估表明:在Robomimic、RLBench、PushT和CortexBench四个基准上,不同视觉编码器与策略学习方法下,成功率达1.3%–10.2%提升;在真实世界中,于UR5及双臂Franka/UR5平台上,抓取、灵巧操作与双臂任务的性能提升达3.3%–23.3%。进一步验证了H2R在跨机体泛化及与视觉-语言-动作模型兼容性方面的潜力。结果表明,H2R通过缓解人机域间视觉差异,增强了机器人策略的泛化能力。

原文摘要 · Abstract (English)

Large-scale pre-training using egocentric human videos has proven effective for robot learning. However, the models pre-trained on such data can be suboptimal for robot learning due to the significant visual gap between human hands and those of different robots. To remedy this, we propose H2R, a human-to-robot data augmentation pipeline that converts egocentric human videos into robot-centric visual data. H2R estimates human hand pose from videos, retargets the motion to simulated robotic arms, removes human limbs via segmentation and inpainting, and composites rendered robot embodiments into the original frames with camera-aligned geometry. This process explicitly bridges the visual gap between human and robot embodiments during pre-training. We apply H2R to augment large-scale egocentric human video datasets such as Ego4D and SSv2. To verify the effectiveness of the augmentation pipeline, we introduce a CLIP-based image-text similarity metric that quantitatively evaluates the semantic fidelity of robot-rendered frames to the original human actions. We evaluate H2R through comprehensive experiments in both simulation and real-world settings. In simulation, H2R consistently improves downstream success rates across four benchmark suites-Robomimic, RLBench, PushT, and CortexBench-yielding gains of 1.3%-10.2% across different visual encoders and policy learning methods. In real-world experiments, H2R improves performance on UR5 and dual-arm Franka/UR5 manipulation platforms, achieving 3.3%-23.3% success rate gains across gripper-based, dexterous, and bimanual tasks. We further demonstrate the potential of H2R in cross-embodiment generalization and its compatibility with vision-language-action models. These results indicate that H2R improves the generalization ability of robotic policies by mitigating the visual discrepancies between human and robot domains.

机器人预训练数据增强视觉-语言-动作跨机体泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。