提出GigaWorld-1世界模型,解决机器人策略评估难问题。
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

- 用真实机器人数据构建评测基准WMBench,支持多维度对比
- 发现长时序动作一致性比视觉逼真度更影响评估可靠性
- 适合研究机器人基础模型评估的学者与工程团队使用
机器人基础模型的评估仍是关键瓶颈;与可高效通过数字基准评估的大语言模型不同,机器人策略依赖缓慢且昂贵的真实世界推演,受限于硬件和人工监督。这推动了以世界模型作为替代评估器的研究。然而,决定世界模型可靠性的核心特性仍不明确。本文系统研究了用于机器人策略评估的世界模型,提出了WMBench基准,基于真实机器人遥操作数据与匹配的策略推演,覆盖多样化的操控任务,支持对模型家族、动作编码、推演时长和评估指标的受控比较。利用WMBench,我们分析了7个视频世界模型、4种动作表示方案及超过32.4万次模拟策略推演,结合来自CVPR 2026 GigaBrain挑战赛的社区提交、定制合成轨迹以及超过1.2万小时的训练视频。实验得出三大核心洞察:评估质量主要取决于长时序、动作忠实的推演一致性,而非短期视觉真实性;预训练收益不仅来自数据规模,更源于通用世界知识与机器人可控性之间的平衡;架构选择如动作编码、记忆设计及评估导向的后训练显著影响与真实机器人行为的一致性。基于此,我们制定实用设计路线并实现为专为策略评估优化的GigaWorld-1,完整开源代码、模型、数据集与工具包,推动具身基础模型可扩展评估研究。
原文摘要 · Abstract (English)
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in \textit{GigaWorld-1}, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。