arXiv:2607.01060cs.RO2026-07被引 1

用快速可靠的神经模拟器,高效评估通用机器人策略。

RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation

论文配图:RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
图 1 · 摘自论文原文
  • 结合自回归视频模型与任务进展感知的评分机制。
  • 长时序推理误差小,真实世界相关性达0.989(皮尔逊)。
  • 适合大规模机器人策略评估,尤其关注效率与可靠性。

视频世界模型正成为评估通用机器人策略的可扩展替代方案,避免了真实部署中的物理限制和工程负担。然而,使用视频世界模型评估策略仍具挑战性,因模型误差导致生成轨迹不可靠,且推理速度慢限制了大规模吞吐。本文提出RoboWorld,一个自动化评估流程,结合快速自回归视频世界模型与任务进展感知的视觉-语言评分模型。为实现可靠长时序自回归世界模型推演,提出步进强制(Step Forcing),通过锚定与单步自前向上下文融合,减少训练-测试不匹配,同时保持动作-观测动态。上述组件使RoboWorld在多种任务与环境中与真实机器人评估高度一致,皮尔逊相关系数r = 0.989,斯皮尔曼等级相关系数ρ = 0.970。

原文摘要 · Abstract (English)

Video world models are emerging as a scalable alternative for evaluating generalist robot policies, bypassing the physical constraints and engineering burdens of real-world deployment. However, evaluating policies with video world models remains challenging, as world-model errors can make generated rollouts unreliable and slow inference limits large-scale throughput. We introduce RoboWorld, an automated evaluation pipeline that pairs a fast autoregressive video world model with a task-progress-aware vision-language model scoring. To enable reliable long-horizon autoregressive world-model rollouts, we propose Step Forcing, which combines anchored and one-step self-forwarded contexts to reduce train-test mismatch while preserving action-observation dynamics. Together, these components enable RoboWorld to align strongly with real-world robot evaluation across tasks and environments, achieving Pearson's r = 0.989 and Spearman's $ρ$ = 0.970.

机器人世界模型仿真评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。