arXiv:2606.00054cs.ROcs.AI2026-06中稿 · IJCAI综述被引 7

用人类视频训练机器人操控模型,突破传统数据成本高、泛化差的瓶颈。

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

论文配图:From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data
图 1 · 摘自论文原文
  • 从人类视频中提取动作线索,分四类方法构建视觉-语言-动作模型
  • 解决视频无标注、机器人与人动作差异大等关键挑战
  • 适合研究机器人泛化控制与多模态学习的学者参考

通用具身控制的进展依赖于大规模视觉-语言-动作(VLA)模型的预训练。然而,现有方法多依赖大量机器人示范数据,获取成本高且与特定机器人绑定。相比之下,人类视频资源丰富,蕴含丰富的语义和物理交互信息,可为真实世界操作提供多样线索。但因体感差异及任务对齐标注缺失,其直接用于VLA模型仍具挑战。本综述系统梳理了如何将人类视频转化为有效知识以训练VLA模型的方法,按动作信息提取方式分为四类:(i) 潜在动作表示,编码帧间变化;(ii) 预测性世界模型,预测未来帧;(iii) 显式2D监督,提取图像平面线索;(iv) 显式3D重建,恢复几何或运动。此外,提出三大开放挑战:将非结构化视频转化为可训练的训练片段,将视频衍生监督映射到机器人可执行动作(应对体感与视角差异),设计更贴近实际部署性能的评估协议。相关论文与资源列表见 https://github.com/AaronFengZY/HumanCentricToVLA-Survey。

原文摘要 · Abstract (English)

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly to obtain and tightly coupled to specific embodiments. Human videos, by contrast, are abundant and capture rich interactions, providing diverse semantic and physical cues for real-world manipulation. Yet, embodiment differences and the frequent absence of task-aligned annotations make their direct use in VLA models challenging. This survey provides a unified view of how human videos are transformed into effective knowledge for VLA models. We categorize existing approaches into four classes based on the action-related information they derive: (i) latent action representations that encode inter-frame changes; (ii) predictive world models that forecast future frames; (iii) explicit 2D supervision that extracts image-plane cues; and (iv) explicit 3D reconstruction that recovers geometry or motion. Beyond this taxonomy, we highlight three key open challenges in this area: structuring unstructured videos into training-ready episodes, grounding video-derived supervision into robot-executable actions under embodiment and viewpoint heterogeneity, and designing evaluation protocols that better predict real-world deployment performance and transfer efficiency, thereby informing future research directions. A curated list of papers and resources is available at https://github.com/AaronFengZY/HumanCentricToVLA-Survey.

视觉-语言-动作人类视频机器人操控多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。