从互联网视频中学习连续运动表示,提升机器人零样本泛化能力
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- 引入时序差分与时间对比学习,强化运动特征并抑制背景干扰
- 在模拟和真实场景中,联合训练策略性能超越现有方法
- 适合需要海量无标注视频数据的机器人动作学习研究者
从互联网视频中无监督学习潜在运动表示对机器人学习至关重要。现有离散方法通常通过小码本的向量量化缓解因提取过多静态背景导致的捷径学习问题,但存在信息丢失,难以捕捉复杂精细的动态。此外,离散潜在运动分布与连续机器人动作之间存在固有差距,阻碍统一策略的联合学习。我们提出CoMo,旨在从大规模互联网视频中学习更精确的连续潜在运动。CoMo采用早期时序差分(Td)机制增加捷径学习难度,并显式增强运动线索;为进一步确保潜在运动聚焦有意义前景,提出时间对比学习(Tcl)方案:正样本使用小未来帧偏移构建,负样本通过直接反转时间方向生成。所提Td与Tcl协同作用,有效使潜在运动更关注前景并强化运动信号。关键的是,CoMo展现出强零样本泛化能力,可为未见视频生成有效伪动作标签。大量仿真与真实世界实验表明,与CoMo伪动作标签联合训练的策略在扩散模型与自回归架构下均表现更优。
原文摘要 · Abstract (English)
Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting excessive static backgrounds through vector quantization with a small codebook size. However, they suffer from information loss and struggle to capture more complex and fine-grained dynamics. Moreover, there is an inherent gap between the distribution of discrete latent motion and continuous robot action, which hinders the joint learning of a unified policy. We propose CoMo, which aims to learn more precise continuous latent motion from internet-scale videos. CoMo employs an early temporal difference (Td) mechanism to increase the shortcut learning difficulty and explicitly enhance motion cues. Additionally, to ensure latent motion better captures meaningful foregrounds, we further propose a temporal contrastive learning (Tcl) scheme. Specifically, positive pairs are constructed with a small future frame temporal offset, while negative pairs are formed by directly reversing the temporal direction. The proposed Td and Tcl work synergistically and effectively ensure that the latent motion focuses better on the foreground and reinforces motion cues. Critically, CoMo exhibits strong zeroshot generalization, enabling it to generate effective pseudo action labels for unseen videos. Extensive simulated and real-world experiments show that policies co-trained with CoMo pseudo action labels achieve superior performance with both diffusion and auto-regressive architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。