让游戏世界状态可直接控制,提升长期生成一致性与可控性。
Marionette: Predicting World States, Rendering Geometry, Painting Appearance

- 显式建模276维3D世界状态,分离几何与外观生成
- 长时序下角色间距从21.2米降至5米,穿透率降66%
- 适合需要精确控制与长期一致性的交互式游戏生成任务
交互式游戏世界模型通常在像素或隐空间中自回归生成视觉观测,迫使姿态、几何和遮挡等结构属性由同一生成序列隐式维持。长时间推演下,这些隐含世界属性的误差累积,导致一致性和可控性变差。本文提出Marionette,一种面向带关节角色的交互式游戏世界模型。首先,双阶段自回归动态模型预测一个显式且可解释的276维3D世界状态,包含多实体关节骨架、度量根轨迹与旋转。其次,零参数图形桥将预测状态转换为姿态控制视频,以闭式计算世界空间几何与遮挡。第三,控制条件下的视频扩散观测模型从结构化控制中合成逼真RGB图像。实验表明:1)强制不匹配动作流使根对齐关节误差降低31%(48个保留片段);2)长时序行为由状态决定,可通过状态规则修复——未干预时两角色漂移至21.2米远(记录值约5米),三分之一帧出现地面穿透;引入地形碰撞器与间距限制后,穿透减少66%,角色保持互动,无需修改观测模型。通过预测状态生成外观,感知保真度无损失(FVD 831 vs. 记录值799)。
原文摘要 · Abstract (English)
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two-stage autoregressive dynamics model predicts an explicit and interpretable 276-dimensional 3D world state comprising multi-entity articulated skeletons, metric root trajectories, and rotations. Second, a zero-parameter graphics bridge converts the predicted state into pose-control videos, computing world-space geometry and occlusion in closed form. Third, a control-conditioned video-diffusion observation model synthesizes photorealistic RGB observations from the resulting structured controls. Our experiments establish two properties of Marionette. First, the predicted world state is directly controllable. Forcing a mismatched action stream changes root-aligned joint error by 31% across 48 held-out segments. Second, long-horizon behaviour is determined in the state, and can be repaired there. Left free, the two generated characters drift to 21.2 m apart (recorded sessions stay near 5 m) and a third of frames show ground penetration. Two rules imposed on the explicit state, a terrain collider and a separation cap, cut penetration by 66% and keep the pair engaged, with no change to the observation model. Routing appearance through the predicted state costs no fidelity we can detect, at an FVD of 831 against 799 for recorded pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。