用仿真环境+真人反馈,实现机器人策略的大规模高效评测。
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
- 将真实机器人视频转为仿真环境,自动构建数字孪生测试场。
- 结合AI评分与众包偏好判断,实现可复现的规模化评估。
- 支持多维度环境扰动,检验策略在复杂变化下的泛化能力。
追求通用机器人代理(能执行多样任务、适应多样环境)需要严格且可扩展的评估手段。然而,真实世界测试机器人策略仍受限于人力密集、速度慢、安全性差及难以复现等问题。随着策略范围和复杂度增加,这些挑战愈发严峻,因机器人任务成功与否常依赖对执行质量的细腻人类判断。本文提出 RobotArena Infinity,一个通过将视觉-语言-动作(VLA)评估迁移至大规模仿真环境并融合在线人类反馈的新基准框架。利用视觉-语言模型、2D到3D生成建模及可微渲染技术,该方法可自动将广泛使用的机器人数据集中的视频示范转化为仿真对应版本。在这些数字孪生环境中,采用基于视觉-语言模型的自动化评分与可扩展的众包偏好判断相结合的方式评估VLA策略,将人类参与从繁琐的场景搭建、重置和安全监督转变为轻量级的偏好比较。为衡量鲁棒性,系统性地在纹理、物体位置等多个维度扰动仿真环境,以在可控变化下压力测试策略的泛化能力。最终形成一个持续演进、可复现、可扩展的基准,用于评估真实世界训练的机器人操作策略,填补了当前机器人领域的一项关键空白。
原文摘要 · Abstract (English)
The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to reproduce. As policies expand in scope and complexity, these barriers only intensify, since defining "success" in robotics often hinges on nuanced human judgments of execution quality. We introduce RobotArena Infinity, a new benchmarking framework that overcomes these challenges by shifting vision-language-action (VLA) evaluation into large-scale simulated environments augmented with online human feedback. Leveraging advances in vision-language models, 2D-to-3D generative modeling, and differentiable rendering, our approach automatically converts video demonstrations from widely used robot datasets into simulated counterparts. Within these digital twins, we assess VLA policies using both automated vision-language-model-guided scoring and scalable human preference judgments collected from crowdworkers, transforming human involvement from tedious scene setup, resetting, and safety supervision into lightweight preference comparisons. To measure robustness, we systematically perturb simulated environments along multiple axes, including textures and object placements, stress-testing policy generalization under controlled variation. The result is a continuously evolving, reproducible, and scalable benchmark for real-world-trained robot manipulation policies, addressing a critical missing capability in today's robotics landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。