构建大规模真实机器人评估系统,测试视觉语言模型的控制能力。
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
- 搭建在线评估平台RoboChallenge,支持多模型多任务实机测试
- 基于Table30基准对数十个前沿视觉语言模型进行实机评测
- 解决真实场景下算法可复现与规模化测试难题,适合机器人研发者参考
真实机器人的测试对机器人控制算法至关重要。在基于学习的算法,尤其是视觉语言模型(VLA)中,大规模评估——即在大量任务上测试众多模型——的需求日益迫切。然而,实现这一目标极具挑战性,尤其是在可扩展性和可复现性方面。本文介绍了我们构建RoboChallenge在线评估系统的方案,并利用初始基准Table30对近期最先进的VLA模型进行了调查评估。
原文摘要 · Abstract (English)
Testing on real machines is indispensable for robotic control algorithms. In the context of learning-based algorithms, especially VLA models, demand for large-scale evaluation, i.e. testing a large number of models on a large number of tasks, is becoming increasingly urgent. However, doing this right is highly non-trivial, especially when scalability and reproducibility is taken into account. In this report, we describe our methodology for constructing RoboChallenge, an online evaluation system to test robotic control algorithms, and our survey of recent state-of-the-art VLA models using our initial benchmark Table30.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。