arXiv:2609.08209cs.RO2026-09

为机器人看视频学技能提供统一评测基准,推动该领域可比性研究。

Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation

论文配图:Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation
图 1 · 摘自论文原文
  • 构建包含真实演示视频与仿真轨迹的统一评测集
  • 覆盖10个操作任务,测试模型对干扰和长序列任务的鲁棒性
  • 涵盖7种前沿算法,揭示当前模型在复杂任务上的不足

观察学习(LfO)是机器人模仿人类与动物社会学习能力的核心技术,尤其适用于数据稀缺的机器人领域。尽管已有研究展示从人类视频中学习操作技能的潜力,但因方法假设、硬件配置和环境设置差异大,难以有效比较进展。为此,我们提出RoboReel:一个统一的评测基准,包含10个操作任务的真实人类示范视频、模拟机器人轨迹及评估环境。通过四个测试套件,评估模型在视觉干扰下的鲁棒性与长时序任务完成能力。基准覆盖多种类别的观察学习模型,并评估了七种先进算法(包括基于视觉-语言-动作模型的变体)。分析表明,长周期任务与容差低的任务仍是当前模型的主要挑战。

原文摘要 · Abstract (English)

Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models' performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io

机器人学习观察学习评测基准视觉导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。