arXiv:2601.11421cs.ROcs.AI2026-01被引 4

构建100个精细化任务,评估机器人智能体真实能力

The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents

  • 基于人机交互原理设计100项任务,覆盖多样化行为
  • 实测验证任务可行且能有效区分现有视觉语言模型表现
  • 适合评估机器人学习方法的综合能力,推动数据集多样性

随着机器人学习与模仿学习的快速发展,大量数据集和方法涌现,但现有数据集的任务设计常缺乏系统性原则。这引发关键问题:当前数据集是否真正提升机器人能力?少数通用任务的评估能否准确反映不同方法的性能差异?为此,我们提出首个机器人学习奥运会雏形——大步100(GM-100),包含100个精心设计的任务,涵盖广泛交互与长尾行为,旨在提供多样且具有挑战性的评估基准,全面检验机器人智能体能力,并促进任务设计的复杂性与多样性。任务设计基于对现有任务的系统分析与扩展,结合人类-物体交互基元与物体重用性洞察。我们在多种机器人平台上采集大量轨迹数据,并评估多个基线模型。实验表明,GM-100任务具备可执行性,且足够具挑战性,能有效区分当前视觉-语言-动作模型的表现。数据与代码已公开于https://rhos.ai/research/gm-100。

原文摘要 · Abstract (English)

Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack systematic consideration and principles. This raises important questions: Do the current datasets and task designs truly advance the capabilities of robotic agents? Do evaluations on a few common tasks accurately reflect the differentiated performance of various methods proposed by different teams and evaluated on different tasks? To address these issues, we introduce the Great March 100 (\textbf{GM-100}) as the first step towards a robot learning Olympics. GM-100 consists of 100 carefully designed tasks that cover a wide range of interactions and long-tail behaviors, aiming to provide a diverse and challenging set of tasks to comprehensively evaluate the capabilities of robotic agents and promote diversity and complexity in robot dataset task designs. These tasks are developed through systematic analysis and expansion of existing task designs, combined with insights from human-object interaction primitives and object affordances. We collect a large amount of trajectory data on different robotic platforms and evaluate several baseline models. Experimental results demonstrate that the GM-100 tasks are 1) feasible to execute and 2) sufficiently challenging to effectively differentiate the performance of current VLA models. Our data and code are available at https://rhos.ai/research/gm-100.

机器人学习任务评估数据集设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。