arXiv:2505.19017cs.ROcs.CV2025-05被引 58

用世界模型模拟真实机器人动作,高效评估机械臂策略性能。

WorldEval: World Model as Real-World Robot Policies Evaluator

  • 将机器人动作编码为隐变量,驱动视频生成模型还原真实操作
  • 在12个任务上实现与真实场景90%以上的性能相关性
  • 适合快速验证新策略安全性和跨环境泛化能力

机器人领域已取得通用抓取策略的显著进展,但真实场景中评估这些策略耗时且困难,尤其在任务数量增多或环境变化时。本文证明世界模型可作为可扩展、可复现、可靠的现实机器人策略评估代理。关键挑战在于从世界模型生成准确反映机器人动作的策略视频。我们发现直接输入机器人动作或使用高维编码方法常无法生成符合动作的视频。为此,提出Policy2Vec,一种将视频生成模型转化为遵循隐动作的仿真器的方法。进而构建WorldEval,一个完全在线的自动化评估流水线,能有效排名不同策略及同一策略的各个检查点,并作为安全检测器防止新模型产生危险动作。通过在真实环境中对多种操作策略进行成对评估,我们展示了WorldEval表现与真实场景高度相关。此外,本方法显著优于主流的实转模方法。

原文摘要 · Abstract (English)

The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time-consuming and challenging, particularly as the number of tasks scales and environmental conditions change. In this work, we demonstrate that world models can serve as a scalable, reproducible, and reliable proxy for real-world robot policy evaluation. A key challenge is generating accurate policy videos from world models that faithfully reflect the robot actions. We observe that directly inputting robot actions or using high-dimensional encoding methods often fails to generate action-following videos. To address this, we propose Policy2Vec, a simple yet effective approach to turn a video generation model into a world simulator that follows latent action to generate the robot video. We then introduce WorldEval, an automated pipeline designed to evaluate real-world robot policies entirely online. WorldEval effectively ranks various robot policies and individual checkpoints within a single policy, and functions as a safety detector to prevent dangerous actions by newly developed robot models. Through comprehensive paired evaluations of manipulation policies in real-world environments, we demonstrate a strong correlation between policy performance in WorldEval and real-world scenarios. Furthermore, our method significantly outperforms popular methods such as real-to-sim approach.

机器人世界模型策略评估仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。