用视觉目标差分构建统一动作空间,实现人与机器人的知识迁移。
IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
- 将图像变化压缩为潜在动作,构建跨人机的统一动作空间。
- 可迁移视频中物体运动至其他场景,支持跨模态动作生成。
- 支持自然语言对齐与底层控制,适用于多任务机器人控制。
我们提出图像-目标表征(IGOR),旨在学习一个统一且语义一致的动作空间,覆盖人类与各类机器人。通过将初始图像与目标状态间的视觉变化压缩为潜在动作,IGOR实现了大规模机器人与人类活动数据的知识迁移。该方法可为互联网规模视频数据生成潜在动作标签,支持在多样化任务中训练基础策略模型与世界模型。实验表明:(1) IGOR能为人类与机器人学习到语义一致的动作空间,表征物体物理交互知识;(2) 结合潜在动作模型与世界模型,可将某视频中的物体运动“迁移”至其他视频,包括跨人类与机器人场景;(3) 通过基础策略模型实现潜在动作与自然语言对齐,并与低层控制模型结合,实现有效机器人控制。我们认为IGOR为人类到机器人的知识迁移与控制开辟了新路径。
原文摘要 · Abstract (English)
We introduce Image-GOal Representations (IGOR), aiming to learn a unified, semantically consistent action space across human and various robots. Through this unified latent action space, IGOR enables knowledge transfer among large-scale robot and human activity data. We achieve this by compressing visual changes between an initial image and its goal state into latent actions. IGOR allows us to generate latent action labels for internet-scale video data. This unified latent action space enables the training of foundation policy and world models across a wide variety of tasks performed by both robots and humans. We demonstrate that: (1) IGOR learns a semantically consistent action space for both human and robots, characterizing various possible motions of objects representing the physical interaction knowledge; (2) IGOR can "migrate" the movements of the object in the one video to other videos, even across human and robots, by jointly using the latent action model and world model; (3) IGOR can learn to align latent actions with natural language through the foundation policy model, and integrate latent actions with a low-level policy model to achieve effective robot control. We believe IGOR opens new possibilities for human-to-robot knowledge transfer and control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。