首个闭环世界模型评测平台,验证生成模型能否真正助力智能体决策。
World-in-World: World Models in a Closed-Loop World
- 构建闭环环境评估世界模型的决策能力,突破传统只看画质的局限。
- 发现控制性比视觉质量更关键,且推理时增加算力可显著提升表现。
- 适合研究世界模型、具身智能与强化学习的学者参考。
生成式世界模型(WMs)现已能模拟出高度逼真的虚拟世界,这自然引发一个问题:它们能否赋予具身智能体预测性感知以支持决策?现有研究受限于评价碎片化——多数基准采用开环协议,仅关注视觉质量,未能解决核心问题:世界模型是否真正帮助智能体完成具身任务?为填补这一空白,我们提出「世界中的世界」(World-in-World),首个在闭环环境中评估世界模型的开放平台。该平台提供统一在线规划策略与标准化动作接口,使异构世界模型可用于决策。我们构建了四个闭环环境,严格评估多种世界模型,以任务成功率为核心指标,超越对视觉质量的片面追求;并首次揭示世界模型在具身场景下的数据缩放规律。研究发现三个意外:(1)仅视觉质量高不足以保证任务成功,可控性更为关键;(2)用动作-观测数据进行后训练扩展,效果优于升级预训练视频生成器;(3)增加推理阶段计算资源可大幅提升闭环性能。
原文摘要 · Abstract (English)
Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has been limited by fragmented evaluation: most existing benchmarks adopt open-loop protocols that emphasize visual quality in isolation, leaving the core issue of embodied utility unresolved, i.e., do WMs actually help agents succeed at embodied tasks? To address this gap, we introduce World-in-World, the first open platform that benchmarks WMs in a closed-loop world that mirrors real agent-environment interactions. World-in-World provides a unified online planning strategy and a standardized action API, enabling heterogeneous WMs for decision making. We curate four closed-loop environments that rigorously evaluate diverse WMs, prioritize task success as the primary metric, and move beyond the common focus on visual quality; we also present the first data scaling law for world models in embodied settings. Our study uncovers three surprises: (1) visual quality alone does not guarantee task success, controllability matters more; (2) scaling post-training with action-observation data is more effective than upgrading the pretrained video generators; and (3) allocating more inference-time compute allows WMs to substantially improve closed-loop performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。