REALM构建真实与仿真间强关联的机器人操作基准,评估模型泛化能力。
REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- 通过高保真视觉与控制对齐,建立仿真与真实性能的强相关性。
- 包含15类扰动、7种技能和3500+物体,覆盖复杂泛化场景。
- 揭示当前视觉语言动作模型在真实世界泛化仍存显著短板。
视觉-语言-动作(VLA)模型使机器人能够理解并执行自然语言指令的任务。然而,其在训练环境之外的泛化能力仍难以有效评估,尤其在真实世界中成本高昂。为此,我们提出REALM——一个新型仿真环境与基准,旨在通过高保真视觉与对齐的机器人控制,建立仿真与真实表现之间的强相关性。该环境包含15种扰动因素、7种操作技能和超过3500个物体。我们构建了两个任务集,并评估了π_{0}、π_{0}-FAST及GR00T N1.5 VLA模型,结果表明泛化与鲁棒性仍是开放挑战。更广泛而言,仿真可作为真实世界的可靠代理,系统性揭示并量化VLA的弱点与失效模式。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they were trained on, which is presently difficult and expensive to evaluate in the real-world. To address this gap, we present REALM, a new simulation environment and benchmark designed to evaluate the generalization capabilities of VLA models, with a specific emphasis on establishing a strong correlation between simulated and real-world performance through high-fidelity visuals and aligned robot control. Our environment offers a suite of 15 perturbation factors, 7 manipulation skills, and more than 3,500 objects. Finally, we establish two task sets that form our benchmark and evaluate the π_{0}, π_{0}-FAST, and GR00T N1.5 VLA models, showing that generalization and robustness remain an open challenge. More broadly, we also show that simulation gives us a valuable proxy for the real-world and allows us to systematically probe for and quantify the weaknesses and failure modes of VLAs. Project page: https://martin-sedlacek.com/realm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。