评测机器人能否理解示范意图而非仅模仿动作。
The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

- 设计四层阶梯式基准,逐步拉开演示与机器人环境差距。
- 9个模型在第三层表现骤降,说明理解意图是关键瓶颈。
- 用少量真人-机器人配对数据微调,性能显著提升。
人类模仿关注意图:观察示范后推断目标,并利用手边工具和场景完成任务。当前机器人策略则直接从视觉输入和语言指令学习观测到动作的映射,未显式推断示范任务。因此,基于人类视频的学习仍局限于轨迹级:模型可在近似场景中复现动作,但难以实现意图层面的模仿。为此,我们提出《模仿者游戏》(The Imitator Game),一个四级基准(L0-L3),逐级扩大演示与机器人场景的差异,揭示轨迹复现失效、任务理解成为必需的临界点。配套构建了IG-10K数据集,目前最大且跨四个层级的环境对齐真人-机器人配对数据集(20,000+配对剧集,50+任务,6个领域),并在真实与仿真环境中均实现覆盖。同时推出开放平台Imitator Arena,支持盲法人机评估。在九个顶尖模型上测试发现,性能从L0到L2稳定,但在L3急剧下降,表明功能替代——通过不同物体属性达成相同目标——是意图级模仿的决定性障碍。以人类视频为条件的模型优于以描述文本为条件的模型,但所有模型在未见任务上的零样本成功率均低于13%;仅用10组真人-机器人配对数据微调预训练模型,即可获得显著提升,且增益随预训练规模增大而增长。
原文摘要 · Abstract (English)
Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution - achieving the same intent through a different object affordance - as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only $10$ paired human-robot demonstrations yields large gains that grow with pretraining scale. The project website and access to Imitator Arena are available at https://imitator-game.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。