提出可量化机器人操作效率与安全性的系统化评估框架
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
- 设计8个双臂任务,控制变量并提供超3000次专家示范
- 引入效率、协调性、稳定性等标准化指标,定位失败环节
- 适合研究机器人策略评估与强化学习的开发者使用
我们提出RoboEval,一个结构化的机器人操作评估框架与基准。现有评估常将性能简化为成功次数,掩盖执行质量差异并模糊失败模式。RoboEval包含8个双臂任务,具有系统性控制变量,提供三千余次专家演示,并配备模块化仿真平台以支持可复现实验。所有任务均配置标准化指标,量化效率、协调性与安全/稳定性,同时追踪阶段性进展并定位失败原因。通过在先进视觉-运动策略上的大量实验,验证了这些指标在不同条件下的稳定性、对相似成功率策略的区分能力,以及与任务成功之间的相关性。
原文摘要 · Abstract (English)
We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled behavioral and outcome metrics. Existing evaluations often collapse performance into outcome counts, masking differences in execution quality and obscuring failure structure. RoboEval provides eight bimanual tasks with systematically controlled variations, more than three thousand expert demonstrations, and a modular simulation platform for reproducible experimentation. All tasks are instrumented with standardized metrics that quantify efficiency, coordination, and safety/stability, as well as outcome measures that trace stagewise progress and localize failure modes. Through extensive experiments with state-of-the-art visuomotor policies, we validate these metrics by analyzing their stability under variation, discriminative power across policies with similar success rates, and correlation with task success. Project Page: https://robo-eval.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。