arXiv:2605.20811cs.RO2026-05被引 1

让机器人跨形态模仿人类动作,无需对齐具体动作

Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation

论文配图:Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation
图 1 · 摘自论文原文
  • 用共享表示空间将视觉演示转为目标机器人的潜在轨迹
  • 在真实任务中实现跨形态模仿,效果接近专用规划器
  • 适合异构机器人迁移学习,仅需视觉示范和自身经验

机器人模仿学习常被视作复现示范动作,但动作本质上依赖具体身体结构。当示范来自不同形态、运动学或动作空间的机器人或人类时,传统方法需共享动作空间、启发式重定向或大规模多身体共训练。本文提出新视角:将示范视为未来目标的隐式描述——目标代理应推断示范者试图达成的状态,而非其执行方式。我们提出Demo-JEPA,一种基于JEPA的世界模型框架,可解耦演示意图与具体执行。该框架将源端视觉示范转化为目标兼容的未来潜在轨迹,置于共享预测表示空间中。目标代理随后利用这些潜在轨迹作为子目标,结合自身学习的前向动态进行规划。由于避免了动作级对应关系,且仅需视觉示范与自身交互经验,该方法支持异构体之间的灵活模仿。在RLBench及真实世界操作任务上的实验表明,Demo-JEPA性能媲美专用领域规划器,并能推广至未见任务与体形配置,而此前方法在此类场景中失败。

原文摘要 · Abstract (English)

Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with different morphology, kinematics, or action spaces, this action-centric view requires shared action spaces, heuristic retargeting, or large-scale multi-embodiment co-training. We instead view demonstrations as implicit specifications of future goals: the target agent should infer what state the demonstrator is trying to realize, rather than how the demonstrator executes it. We propose Demo-JEPA, a cross-embodiment imitation framework that decouples demonstration intent from embodiment-specific execution. Built on a JEPA-based world model, Demo-JEPA translates source visual demonstrations into target-compatible future latent trajectories in a shared predictive representation space. The target agent then uses these latent trajectories as subgoals and realizes them through planning under its own learned forward dynamics. Because Demo-JEPA avoids action-level correspondence and requires only visual demonstrations plus the target agent's own interaction experience, it supports flexible imitation across heterogeneous embodiments. Experiments on RLBench and real-world manipulation tasks show that Demo-JEPA matches specialized in-domain planners and generalizes to unseen tasks and embodiment configurations where prior methods fail.

模仿学习跨形态世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。