arXiv:2606.16993cs.CV2026-06被引 17

一个可交互的通用视频世界模型,支持长时序控制生成。

DreamX-World 1.0: A General-Purpose Interactive World Model

论文配图:DreamX-World 1.0: A General-Purpose Interactive World Model
图 1 · 摘自论文原文
  • 用轻量投影位置编码实现相机感知的注意力机制。
  • 通过自生成长序列训练减少风格漂移,实现16帧/秒实时生成。
  • 适合需要可控长视频生成的研究者与开发者。

DreamX-World 1.0 是一个通用的文本/图像到视频的交互式世界模型,支持在摄影写实、游戏风格和艺术化场景中进行可控的长时序生成,包括相机导航、回访已观察区域和提示驱动事件。其数据引擎结合了相机精确的 Unreal Engine 渲染、动作丰富的游戏录像以及恢复相机几何的真实视频。为实现相机控制,提出 E-PRoPE,一种轻量级投影位置编码,在保留投影相机几何的同时引入相机感知注意力。通过因果强制、类 DMD 蒸馏和长滚动训练,将双向视频生成器转为少步自回归世界模型。在自生成长时序上下文中训练,使模型暴露于自身生成的历史,降低自回归块间的风格与色彩漂移。记忆条件场景持久性通过基于相机几何的检索获取早期视角,残差复用则降低对不完美记忆隐变量的敏感性。事件指令微调实现可组合事件控制,强化学习对齐在蒸馏后恢复相机控制与视觉质量。采用混合精度 DiT 执行、残差复用、75% 剪枝的 VAE 解码和异步流水线并行,该模型在八张 RTX 5090 GPU 上达到最高 16 FPS。在 5 秒基础评估中,相机控制得分为 73.75,综合得分为 84.76,优于 HY-WorldPlay 1.5(80.79)和 LingBot-World(80.45)。

原文摘要 · Abstract (English)

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously observed regions, and promptable events across photorealistic, game-style, and stylized domains. Our data engine combines camera-accurate Unreal Engine rendering, action-rich gameplay recordings, and real-world videos with recovered camera geometry. For camera control, we introduce E-PRoPE, a lightweight variant of projective positional encoding that retains PRoPE's projective camera geometry while applying camera-aware attention to spatially reduced tokens. We convert a bidirectional video generator into a few-step autoregressive world model using causal forcing, DMD-style distillation, and long-rollout training. Training on self-generated long-horizon contexts exposes the model to its own generated history and reduces the style and color drift that accumulates across autoregressive chunks. Memory-Conditioned Scene Persistence retrieves earlier views through camera-geometry-based retrieval, while residual recycling makes the conditioning path less sensitive to imperfect memory latents. Event Instruction Tuning adds composable event control, and reinforcement learning alignment recovers camera control and visual quality after distillation. With mixed-precision DiT execution, residual reuse, 75\%-pruned VAE decoding, and asynchronous pipeline parallelism, DreamX-World 1.0 reaches up to 16\,FPS on eight RTX\,5090 GPUs. On our 5-second basic evaluation, DreamX-World 1.0 achieves a camera-control score of 73.75 and an overall score of 84.76, outperforming HY-WorldPlay 1.5 and LingBot-World in overall score, which achieve 80.79 and 80.45, respectively.

视频生成世界模型交互控制长时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。