arXiv:2503.24278cs.ROcs.AI2025-03被引 54

AutoEval让机器人策略在真实世界自动评估,无需人工干预。

AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World

  • 通过自动成功检测与场景重置,实现无人值守的全天候评估
  • 评估结果与人工标注高度一致,误差极小
  • 支持通用机器人策略评估,适合研究者快速验证新方法

可扩展且可复现的策略评估一直是机器人学习中的长期挑战。真实世界中的评估对人力时间要求高,难以规模化获取可靠结果。随着机器人策略日益通用,所需评估环境也日趋多样,评估瓶颈愈发突出。为使真实世界评估更可行,我们提出 AutoEval 系统,实现通用机器人策略的全自动、全天候评估,仅需极少人工干预。用户通过提交评估任务至队列,类似集群调度系统;AutoEval 在提供自动成功检测与自动场景重置的框架内调度策略执行。实验表明,AutoEval 几乎完全消除了人工参与,支持持续评估,结果与人工地面真值评估高度吻合。为推动通用策略评估在机器人社区的发展,我们公开提供多个基于 BridgeData 机器人配置和 WidowX 机械臂的 AutoEval 场景。未来希望各机构共建分布式评估网络,覆盖更多场景。

原文摘要 · Abstract (English)

Scalable and reproducible policy evaluation has been a long-standing challenge in robot learning. Evaluations are critical to assess progress and build better policies, but evaluation in the real world, especially at a scale that would provide statistically reliable results, is costly in terms of human time and hard to obtain. Evaluation of increasingly generalist robot policies requires an increasingly diverse repertoire of evaluation environments, making the evaluation bottleneck even more pronounced. To make real-world evaluation of robotic policies more practical, we propose AutoEval, a system to autonomously evaluate generalist robot policies around the clock with minimal human intervention. Users interact with AutoEval by submitting evaluation jobs to the AutoEval queue, much like how software jobs are submitted with a cluster scheduling system, and AutoEval will schedule the policies for evaluation within a framework supplying automatic success detection and automatic scene resets. We show that AutoEval can nearly fully eliminate human involvement in the evaluation process, permitting around the clock evaluations, and the evaluation results correspond closely to ground truth evaluations conducted by hand. To facilitate the evaluation of generalist policies in the robotics community, we provide public access to multiple AutoEval scenes in the popular BridgeData robot setup with WidowX robot arms. In the future, we hope that AutoEval scenes can be set up across institutions to form a diverse and distributed evaluation network.

机器人评估自动测试真实世界通用策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。