arXiv:2506.18123cs.ROcs.LG2025-06被引 75

用众包方式在真实世界评估机器人通用策略,更准更可扩展。

RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

  • 通过分布式众包评估,自由选择任务与环境
  • 600+次真实机器人对比实验,准确排序7种策略
  • 双盲对比提升可信度,适合广泛评测通用机器人

现代通用机器人策略的全面、无偏且可比的评估极具挑战:现有基准方法通常依赖固定任务、环境或集中式‘机器人竞赛’,难以扩展至多样任务和环境。本文提出RoboArena,一种面向真实世界的可扩展通用策略评估新方法。不强制统一任务或地点,而是通过分布式的评估网络,让各机构自由选择评估内容,但要求对策略对进行双盲比较。通过聚合跨多样任务与环境的偏好反馈,实现策略排名。我们在七所高校的DROID机器人平台上部署该系统,完成超过600次真实机器人成对评估,覆盖7种通用策略。结果表明,该众包方法在准确性上优于传统集中式评估,同时更具可扩展性、鲁棒性和可信度。我们已向社区开放评估网络,旨在推动通用机器人策略的更广泛可比性。

原文摘要 · Abstract (English)

Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standardization, either by specifying fixed evaluation tasks and environments, or by hosting centralized ''robot challenges'', and do not readily scale to evaluating generalist policies across a broad range of tasks and environments. In this work, we propose RoboArena, a new approach for scalable evaluation of generalist robot policies in the real world. Instead of standardizing evaluations around fixed tasks, environments, or locations, we propose to crowd-source evaluations across a distributed network of evaluators. Importantly, evaluators can freely choose the tasks and environments they evaluate on, enabling easy scaling of diversity, but they are required to perform double-blind evaluations over pairs of policies. Then, by aggregating preference feedback from pairwise comparisons across diverse tasks and environments, we can derive a ranking of policies. We instantiate our approach across a network of evaluators at seven academic institutions using the DROID robot platform. Through more than 600 pairwise real-robot evaluation episodes across seven generalist policies, we demonstrate that our crowd-sourced approach can more accurately rank the performance of existing generalist policies than conventional, centralized evaluation approaches, while being more scalable, resilient, and trustworthy. We open our evaluation network to the community and hope that it can enable more accessible comparisons of generalist robot policies.

机器人评估众包通用策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。