让机器人从人类视频中学习动作,解决视觉与动作不匹配问题。
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

- 用少量配对数据做桥梁,大量单向视频做监督,对齐人机视觉与动作表征。
- 在CALVIN数据集上平均任务长度达4.481,真实场景成功率提升7.1%。
- 适合做机器人模仿学习、跨模态动作建模的研究者和开发者。
从大规模人类视频中学习通用的视觉-语言-动作(VLA)模型具有前景,但受制于视觉观察和可执行动作的跨本体差异。尽管隐式动作模型通过学习动作抽象减少了动作执行差距,仍依赖视觉特征,导致人机视觉表征不一致,引发策略输入不一致并产生依赖领域的隐式动作,阻碍与人类视频的有效联合训练。为此,我们提出HARP框架,实现更有效的从人类视频出发的VLA预训练。具体而言,HARP利用少量成对的人机演示作为跨本体桥梁,以及大量非配对的人类和机器人视频作为可扩展的动力学监督数据源。它训练一个适应机器人的视觉编码器和隐式动作模型,结合以操作为中心的辅助线索和源相对的成对判别对齐损失,使机器人表征向人类语义对齐,同时保持成对级别的判别能力。所学习到的对齐视觉编码器和隐式动作模型为VLA风格策略学习提供统一的视觉与动作表示,其中人类和机器人视频提供视觉-语言到隐式动作的监督,轻量级机器人动作头将隐式动作转化为可执行命令。在特征可视化、仿真和真实世界操作实验中,验证了更好的人机对齐性和下游策略性能,于CALVIN ABC→D任务上达到4.481的平均长度,并在真实世界相比最强基线提升7.1%的成功率。
原文摘要 · Abstract (English)
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent action models reduce the action execution gap by learning action abstractions, they still rely on visual features. Thus, misaligned human and robot visual representations can lead to inconsistencies in policy inputs and induce domain-dependent latent actions, hindering effective co-training with human videos. To address this, we propose HARP, a human-robot aligned representation learning framework for more effective VLA pretraining from human videos. Specifically, HARP uses limited paired human-robot demonstrations as cross-embodiment bridges and abundant unpaired human and robot videos as a scalable dynamics supervision data source. It trains a robot-adapted visual encoder and a latent action model with manipulation-centric auxiliary cues and a source-relative pair-discriminative alignment loss, which adapts robot representations toward human semantics while preserving pair-level discrimination. The learned aligned vision encoder and latent action model provide a unified vision and action representation for VLA-style policy learning, where human and robot videos provide vision-language-to-latent-action supervision and a lightweight robot action head grounds latent actions into executable commands. Experiments on feature visualization, simulation, and realworld manipulation show improved human-robot alignment and downstream policy performance, achieving 4.481 average length on CALVIN ABC$\rightarrow$D and a 7.1\% realworld success rate gain over the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。