从人类视频中无监督学习可迁移的机器人动作表征。
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
- 通过对比解耦机制分离动作动态与视觉内容。
- 仅用人类视频预训练,超越真实机器人轨迹表现。
- 适合需要海量数据但难获取真实操作的场景。
视觉-语言-动作(VLA)模型通过大规模机器人远程操控数据集预训练实现初步泛化。然而,全面覆盖多样任务与环境的数据集获取成本极高且难以扩展。相比之下,人类示范视频提供了丰富且可扩展的多样化场景与操作行为,但缺乏显式动作监督,限制了直接使用。先前工作采用基于VQ-VAE的框架从人类视频中无监督学习潜在动作。然而,由于训练目标主要聚焦于重建视觉外观而非捕捉帧间动态,所学表征往往依赖虚假视觉线索,导致捷径学习和纠缠的潜在表示,阻碍可迁移性。为此,我们提出ConLA,一种从人类视频中无监督预训练机器人策略的框架。ConLA引入对比解耦机制,利用动作类别先验和时序线索,将运动动态与视觉内容分离,有效缓解捷径学习。大量实验表明,ConLA在多种基准上表现优异。值得注意的是,仅在人类视频上预训练,本方法首次超越真实机器人轨迹预训练性能,凸显其提取纯净、语义一致的潜在动作表征的能力,适用于可扩展的机器人学习。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models achieve preliminary generalization through pretraining on large scale robot teleoperation datasets. However, acquiring datasets that comprehensively cover diverse tasks and environments is extremely costly and difficult to scale. In contrast, human demonstration videos offer a rich and scalable source of diverse scenes and manipulation behaviors, yet their lack of explicit action supervision hinders direct utilization. Prior work leverages VQ-VAE based frameworks to learn latent actions from human videos in an unsupervised manner. Nevertheless, since the training objective primarily focuses on reconstructing visual appearances rather than capturing inter-frame dynamics, the learned representations tend to rely on spurious visual cues, leading to shortcut learning and entangled latent representations that hinder transferability. To address this, we propose ConLA, an unsupervised pretraining framework for learning robotic policies from human videos. ConLA introduces a contrastive disentanglement mechanism that leverages action category priors and temporal cues to isolate motion dynamics from visual content, effectively mitigating shortcut learning. Extensive experiments show that ConLA achieves strong performance across diverse benchmarks. Notably, by pretraining solely on human videos, our method for the first time surpasses the performance obtained with real robot trajectory pretraining, highlighting its ability to extract pure and semantically consistent latent action representations for scalable robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。