用视觉视频提取机器人任务的共同终点规律,提升无动作演示下的操作成功率。
CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation

- 从无动作视频中提炼任务终点的共性结构特征作为目标
- 在真实场景中提升成功率最高达17.0个百分点
- 适合缺乏机器人示范数据但有大量视频的场景
机器人模仿学习通常依赖昂贵的机器人示范,而大量无动作的视觉示范(如人类视频)因缺乏可执行动作和具身差异难以利用。本文提出CORE框架,从视觉示范中提取共性终点规律(Common Outcome Regularities)。核心观察是:同一任务的成功轨迹虽多样,但其终点常具有稳定的物体构型、空间关系和接触约束。CORE首先通过对比和辅助时序目标训练终点编码器,再将成功终点嵌入聚类为视觉目标原型,并将其作为全局目标条件注入机器人策略。相比语言指令,视觉目标原型提供了更具体的几何与物理约束。在Meta-World、RoboTwin 2.0和真实世界操作任务中,相较基线策略,成功率分别提升最多3.9、11.1和17.0个百分点,且优于文本条件变体。项目与代码已公开。
原文摘要 · Abstract (English)
Robot imitation learning often relies on costly robot demonstrations, while abundant action-free visual demonstrations, such as human videos, are difficult to use because they lack robot-executable actions and suffer from embodiment gaps. We propose CORE, a policy learning framework that extracts Common Outcome Regularities (CORE) from visual demonstrations. Rather than transferring explicit actions across embodiments, CORE exploits a key observation: although successful trajectories for the same task can be diverse, their terminal states often share stable object configurations, spatial relations, and contact constraints. CORE first trains a terminal outcome encoder with contrastive and auxiliary temporal objectives, then aggregates successful terminal embeddings into visual goal prototypes, and finally injects these prototypes as global goal conditions into robot policies. Compared with language instructions, visual goal prototypes provide more concrete geometric and physical constraints for task completion. Across Meta-World, RoboTwin 2.0, and real-world manipulation, CORE improves the average success rate of the corresponding policy backbones by up to +3.9, +11.1, and +17.0 percentage points, respectively, and outperforms text-conditioned variants under the evaluated settings. The project and code are available at https://logssim.github.io/CORE.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。