低成本机械臂集群实现机器人抓取策略的并行真实世界评估
ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

- 用低成本机械臂农场并行测试策略,降低实验成本
- 完成7种策略在12个任务上的评估,共3118次真实运行
- 提供带质量标签的数据,支持后续学习与模型训练
真实世界评估是开发通用机器人抓取策略的瓶颈,每次试运行都需要硬件和人工操作。我们提出ArmnetBench v0.1,基于轻量现场监督的低成本SO-101机械臂集群,端到端验证该臂农场系统。v0.1对比了7种策略在12个任务中的表现,涵盖单臂与双臂配置;每项任务使用50次示范进行训练或微调,共产生2,518次策略试运行和600次参考示范。所有3,118个任务序列均标注三类结果(成功、次优、失败)。策略试运行由人工评分,示范则默认成功。数据可用于下游学习,如奖励建模、预测世界模型及混合质量数据训练策略。排行榜基于统一预算下的初始比较。核心3,118个数据集已发布于LeRobot v3.0和RoboMeter格式。
原文摘要 · Abstract (English)
Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。