用人类视频学技能,让机器人少练多会。
ViSA-Flow: Accelerating Robot Skill Learning via Large-Scale Video Semantic Action Flow
- 从人类操作视频中自动提取语义动作流作为中间表示
- 仅需少量机器人演示即可实现高性能技能迁移
- 特别适合数据稀缺场景,提升机器人学习效率
制约机器人掌握复杂操作技能的核心挑战是收集大规模机器人示范成本过高。相比之下,人类可通过观察他人与环境交互高效学习。为此,我们引入语义动作流作为核心中间表示,捕捉操纵器-物体间关键时空交互,对表面视觉差异具有不变性。提出ViSA-Flow框架,通过自监督方式从大量未标注视频数据中学习该表示。首先,在大规模人-物交互视频中自动提取语义动作流,预训练生成模型以学习稳健的操作结构先验;其次,通过在少量机器人示范上微调该先验,高效适配至目标机器人。在CALVIN基准和真实任务上的大量实验表明,ViSA-Flow在低数据条件下达到当前最优性能,有效实现了从人类视频观察到机器人执行的知识迁移。视频展示见https://visaflow-web.github.io/ViSAFLOW。
原文摘要 · Abstract (English)
One of the central challenges preventing robots from acquiring complex manipulation skills is the prohibitive cost of collecting large-scale robot demonstrations. In contrast, humans are able to learn efficiently by watching others interact with their environment. To bridge this gap, we introduce semantic action flow as a core intermediate representation capturing the essential spatio-temporal manipulator-object interactions, invariant to superficial visual differences. We present ViSA-Flow, a framework that learns this representation self-supervised from unlabeled large-scale video data. First, a generative model is pre-trained on semantic action flows automatically extracted from large-scale human-object interaction video data, learning a robust prior over manipulation structure. Second, this prior is efficiently adapted to a target robot by fine-tuning on a small set of robot demonstrations processed through the same semantic abstraction pipeline. We demonstrate through extensive experiments on the CALVIN benchmark and real-world tasks that ViSA-Flow achieves state-of-the-art performance, particularly in low-data regimes, outperforming prior methods by effectively transferring knowledge from human video observation to robotic execution. Videos are available at https://visaflow-web.github.io/ViSAFLOW.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。