arXiv:2604.27621cs.ROcs.CV2026-04综述被引 4

用人类视频教机器人学技能,突破数据瓶颈

Robot Learning from Human Videos: A Survey

论文配图:Robot Learning from Human Videos: A Survey
图 1 · 摘自论文原文
  • 从人类动作视频中提取技能,实现机器人被动学习
  • 构建任务/观测/动作三类迁移路径,系统梳理技术体系
  • 梳理主流数据集与生成方法,助力通用机器人发展

制约具身智能与机器人发展的关键瓶颈在于机器人数据的规模化。近年来,利用人类活动视频学习机器人操作技能的研究受到广泛关注,得益于海量的人类行为视频和计算机视觉的进步。该方向有望让机器人从丰富且易获取的人类示范中被动习得技能,显著推动通用机器人系统的可扩展学习。本文综述了基于人类视频的机器人学习技术,聚焦技能迁移与数据基础。首先回顾机器人策略学习基础,接着介绍人类视频的整合接口;随后提出分层分类体系,涵盖任务、观测、动作导向的技能迁移路径,并分析其与不同数据配置和学习范式的耦合关系。此外,系统调研常用人类视频数据集与视频生成方法,揭示数据集发展与应用的大规模统计趋势。最后,指出该领域固有挑战与局限,并展望未来研究方向。论文列表见:https://github.com/IRMVLab/awesome-robot-learning-from-human-videos。

原文摘要 · Abstract (English)

A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted rapidly growing attention in recent years, driven by the abundance of human activity videos and advances in computer vision. This line of research promises to enable robots to acquire skills passively from the vast and readily available resource of human demonstrations, substantially favoring scalable learning for generalist robotic systems. Therefore, we present this survey to provide a comprehensive and up-to-date review of human-video-based learning techniques in robotics, focusing on both human-robot skill transfer and data foundations. We first review the policy learning foundations in robotics, and then describe the fundamental interfaces to incorporate human videos. Subsequently, we introduce a hierarchical taxonomy of transferring human videos to robot skills, covering task-, observation-, and action-oriented pathways, along with a cross-family analysis of their couplings with different data configurations and learning paradigms. In addition, we investigate the data foundations including widely-used human video datasets and video generation schemes, and provide large-scale statistical trends in dataset development and utilization. Ultimately, we emphasize the challenges and limitations intrinsic to this field, and delineate potential avenues for future research. The paper list of our survey is available at https://github.com/IRMVLab/awesome-robot-learning-from-human-videos.

机器人学习视频理解技能迁移数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。