用真人操作视频生成可跨机器人形态使用的动作数据。
RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning
- 从单目相机视频中重建高精度手物交互轨迹,结合物理约束优化姿态。
- 生成的轨迹在多种机器人上可执行,性能接近遥控操作,支持持续学习。
- 适合做具身智能、模仿学习与多形态机器人训练的研究者使用。
我们提出Robowheel,一个将真人手物交互(HOI)视频转化为跨形态机器人学习可用监督信号的数据引擎。基于单目RGB或RGB-D输入,通过高精度HOI重建,并利用强化学习优化器在接触与穿透约束下精炼手物相对位姿,确保物理合理性。重构的富含接触信息的轨迹可被重定向至不同形态的机器人,包括简单末端执行器、灵巧手和人形机器人,生成可执行动作与仿真轨迹。为扩大覆盖范围,我们在Isaac Sim中构建了模拟增强框架,采用多样化领域随机化(包括形态、轨迹、物体获取、背景纹理、手势镜像),丰富轨迹与观测分布,同时保持空间关系与物理合理性。整个数据流程实现从视频到重建、重定向、增强数据获取的端到端闭环。我们在主流视觉语言动作(VLA)与模仿学习架构上验证该数据,结果表明其轨迹稳定性与遥控操作相当,且带来相当的持续性能提升。据我们所知,这是首次定量证明HOI模态可作为有效机器人学习监督信号。相比遥控操作,Robowheel轻量高效,仅需单个单目RGB(D)相机即可提取通用、形态无关的动作表征,灵活适配各类机器人形态。我们还构建了一个大规模多模态数据集,融合多相机采集、单目视频与公开HOI语料库,用于训练与评估具身模型。
原文摘要 · Abstract (English)
We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular RGB or RGB-D inputs, we perform high precision HOI reconstruction and enforce physical plausibility via a reinforcement learning (RL) optimizer that refines hand object relative poses under contact and penetration constraints. The reconstructed, contact rich trajectories are then retargeted to cross-embodiments, robot arms with simple end effectors, dexterous hands, and humanoids, yielding executable actions and rollouts. To scale coverage, we build a simulation-augmented framework on Isaac Sim with diverse domain randomization (embodiments, trajectories, object retrieval, background textures, hand motion mirroring), which enriches the distributions of trajectories and observations while preserving spatial relationships and physical plausibility. The entire data pipeline forms an end to end pipeline from video,reconstruction,retargeting,augmentation data acquisition. We validate the data on mainstream vision language action (VLA) and imitation learning architectures, demonstrating that trajectories produced by our pipeline are as stable as those from teleoperation and yield comparable continual performance gains. To our knowledge, this provides the first quantitative evidence that HOI modalities can serve as effective supervision for robotic learning. Compared with teleoperation, Robowheel is lightweight, a single monocular RGB(D) camera is sufficient to extract a universal, embodiment agnostic motion representation that could be flexibly retargeted across embodiments. We further assemble a large scale multimodal dataset combining multi-camera captures, monocular videos, and public HOI corpora for training and evaluating embodied models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。