构建首个带显式状态标注的大规模动态世界建模数据集,支持动作驱动的长时序生成。
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
- 从真实游戏采集百万级帧,动作与状态解耦,支持结构化建模
- 包含450+语义动作和每帧状态标注,验证模型需保持长期状态一致性
- 适合研究视频生成、强化学习与具身智能的开发者使用
动态系统理论与强化学习将世界演化视为由动作驱动的隐状态动态,视觉观测仅提供状态的部分信息。现有视频世界模型尝试从数据中学习这种动作条件下的动态,但多数数据集缺乏多样且语义明确的动作空间,且动作直接绑定于视觉变化,而非通过底层状态中介。这导致动作与像素级变化纠缠,难以学习结构化世界动态并维持长时序一致性。本文提出WildWorld,一个大规模动作条件世界建模数据集,具有显式状态标注,通过光影视觉化的3A动作角色扮演游戏《怪物猎人:荒野》自动采集。数据集包含超过1.08亿帧,涵盖450多个动作(包括移动、攻击、技能释放),并同步标注每帧的角色骨骼、世界状态、相机位姿和深度图。我们进一步构建WildBench,通过动作跟随与状态对齐评估模型。大量实验揭示了在语义丰富动作建模与长时序状态一致性方面的持续挑战,凸显状态感知视频生成的必要性。
原文摘要 · Abstract (English)
Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this action-conditioned dynamics from data. However, existing datasets rarely match the requirement: they typically lack diverse and semantically meaningful action spaces, and actions are directly tied to visual observations rather than mediated by underlying states. As a result, actions are often entangled with pixel-level changes, making it difficult for models to learn structured world dynamics and maintain consistent evolution over long horizons. In this paper, we propose WildWorld, a large-scale action-conditioned world modeling dataset with explicit state annotations, automatically collected from a photorealistic AAA action role-playing game (Monster Hunter: Wilds). WildWorld contains over 108 million frames and features more than 450 actions, including movement, attacks, and skill casting, together with synchronized per-frame annotations of character skeletons, world states, camera poses, and depth maps. We further derive WildBench to evaluate models through Action Following and State Alignment. Extensive experiments reveal persistent challenges in modeling semantically rich actions and maintaining long-horizon state consistency, highlighting the need for state-aware video generation. The project page is https://shandaai.github.io/wildworld-project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。