arXiv:2606.18955cs.CVcs.RO2026-06中稿 · IROS 2026

用人类第一视角视频训练通用视觉语言动作模型,无需标注动作。

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

论文配图:Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
图 1 · 摘自论文原文
  • 从无标签人类视频中提取动作先验,通过解耦运动与背景构建通用动作码本。
  • 仅用50条轨迹即可在仿真和真实环境实现媲美大规模标注数据的性能。
  • 适合想用免费人类视频数据训练机器人模型的研究者。

训练通用视觉-语言-动作(VLA)模型通常需要大量多样且带有高保真动作标注的机器人数据集。尽管第一人称人类操作视频数量丰富且涵盖大量环境变化,但缺乏动作标签使其难以用于传统训练范式。为此,我们提出一种基于潜在动作的框架,旨在从无标注的人类视频中提取通用动作先验。该架构采用混合解耦式VQ-VAE,通过物理掩码将运动动态与环境背景解耦,从而构建跨体态动作码本。通过在人类视频上预训练码本,视觉语言模型主干学习深层动作意图表示。为适配特定体态,引入意图-感知解耦策略:视觉语言模型预测动作意图,而独立冻结的视觉编码器提供状态特异性特征给动作专家,有效减少动作幻觉。仿真与真实环境结果表明,本方法仅在无标注人类视频上预训练,便能在下游适应中达到与依赖大规模标注数据的顶尖VLA模型相当的性能,仅需50条轨迹即可完成微调。

原文摘要 · Abstract (English)

Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture significant environmental diversity, the absence of action labels makes them difficult to use in conventional training paradigms. To address this, we propose a latent-action-based framework designed to extract general action priors from unlabeled human videos. The architecture features a Hybrid Disentangled VQ-VAE that decouples motion dynamics from environmental backgrounds through physical masks, enabling the construction of a cross-embodiment action codebook. By pre-training on human videos with the codebook, the VLM backbone learns deep representations of action intent. For adaptation to specific embodiments, we introduce an intent-perception decoupling strategy where the VLM predicts the action intent while a separate frozen visual encoder provides state-specific features to the action expert, thereby reducing action hallucinations. Results in simulation and real-world environments show that our method, pre-trained exclusively on unlabeled human videos, performs competitively with state-of-the-art VLA models trained on massive annotated datasets, requiring only 50 trajectories for downstream adaptation.

视觉语言动作无监督训练第一人称视频跨体态迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。