用人类视频教机器人学动作,跨身体形态也能精准模仿。
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
- 从大规模人机视频中无监督学习通用技能表示
- 在仿真与真实环境中实现未见过的人类视频动作迁移
- 适合机器人动作模仿、跨模态技能学习的研究者
模仿是人类学习新任务的核心机制,通过观察和模仿专家来掌握技能。然而,由于人类与机器人在视觉外观和物理能力上的本质差异,将此机制应用于机器人面临巨大挑战。尽管已有方法通过共享场景与任务的跨身体形态数据集来弥合这一差距,但大规模收集人类与机器人对齐的数据仍不现实。本文提出UniSkill,一种新颖框架,仅利用大规模跨身体形态视频数据(无需标签),学习与身体形态无关的技能表示。由此提取的人类视频提示中的技能可有效迁移至仅在机器人数据上训练的策略中。实验在仿真与真实环境均表明,该跨身体形态技能能成功引导机器人选择恰当动作,即使面对未见过的视频提示亦可实现。项目主页:https://kimhanjung.github.io/UniSkill。
原文摘要 · Abstract (English)
Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents significant challenges due to the inherent differences between human and robot embodiments in both their visual appearance and physical capabilities. While previous methods bridge this gap using cross-embodiment datasets with shared scenes and tasks, collecting such aligned data between humans and robots at scale is not trivial. In this paper, we propose UniSkill, a novel framework that learns embodiment-agnostic skill representations from large-scale cross-embodiment video data without any labels, enabling skills extracted from human video prompts to effectively transfer to robot policies trained only on robot data. Our experiments in both simulation and real-world environments show that our cross-embodiment skills successfully guide robots in selecting appropriate actions, even with unseen video prompts. The project website can be found at: https://kimhanjung.github.io/UniSkill.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。