用真人示范生成机器人可执行数据,无需远程操控。
RoboPaint: From Human Demonstration to Any Robot and Any View
- 通过真实-仿真-真实流程,将人手动作转换为机器人动作
- 10项任务中成功率84%,训练模型平均达80%成功
- 适合想低成本获取复杂操作数据的研究者
获取大规模、高保真的机器人示范数据仍是提升灵巧操作中视觉-语言-动作(VLA)模型性能的关键瓶颈。我们提出一种‘真实-仿真-真实’的数据采集与编辑流程,将人类示范转化为无需直接机器人遥控即可执行的环境特定训练数据。在标准化数据采集室内,同步记录三路RGB-D视频、11路RGB视频、29自由度数据手套关节角度及14通道触觉信号。基于这些人类示范,我们引入一种触觉感知的重定向方法,通过几何与力引导优化,将人类手部状态映射至机器人灵巧手状态。随后,将重定向后的机器人轨迹在逼真的Isaac Sim环境中渲染,构建机器人训练数据。真实世界实验表明:(1) 重定向的灵巧手轨迹在10种不同物体操作任务中达到84%成功率;(2) 仅使用我们生成数据训练的VLA策略(Pi0.5)在三个代表性任务(抓取放置、推动、倾倒)上平均成功率可达80%。结论:可通过本真实-仿真-真实数据管道,高效‘绘制’出机器人训练数据。该方法提供了一种可扩展、低成本的替代远程操控方案,对复杂灵巧操作性能损失极小。
原文摘要 · Abstract (English)
Acquiring large-scale, high-fidelity robot demonstration data remains a critical bottleneck for scaling Vision-Language-Action (VLA) models in dexterous manipulation. We propose a Real-Sim-Real data collection and data editing pipeline that transforms human demonstrations into robot-executable, environment-specific training data without direct robot teleoperation. Standardized data collection rooms are built to capture multimodal human demonstrations (synchronized 3 RGB-D videos, 11 RGB videos, 29-DoF glove joint angles, and 14-channel tactile signals). Based on these human demonstrations, we introduce a tactile-aware retargeting method that maps human hand states to robot dex-hand states via geometry and force-guided optimization. Then the retargeted robot trajectories are rendered in a photorealistic Isaac Sim environment to build robot training data. Real world experiments have demonstrated: (1) The retargeted dex-hand trajectories achieve an 84\% success rate across 10 diverse object manipulation tasks. (2) VLA policies (Pi0.5) trained exclusively on our generated data achieve 80\% average success rate on three representative tasks, i.e., pick-and-place, pushing and pouring. To conclude, robot training data can be efficiently "painted" from human demonstrations using our real-sim-real data pipeline. We offer a scalable, cost-effective alternative to teleoperation with minimal performance loss for complex dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。