arXiv:2607.02517cs.CV2026-07被引 3

让视频世界模型具备持久动态记忆,可自由控制物体运动轨迹。

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

论文配图:WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
图 1 · 摘自论文原文
  • 用大模型协调3D轨迹与摄像机运动,分离语义控制与视觉生成
  • 物体离屏后重入仍保持完全一致的视觉身份,支持长时间动态追踪
  • 适合需要精确控制和长期一致性的视频生成任务

我们提出WorldDirector,一种高度可控的视频世界模型框架,支持持久动态对象记忆与任意视角探索。不同于现有世界模型将物理动态与像素渲染耦合,并依赖持续视觉观测维持运动,我们的方法显式解耦语义运动编排与视觉生成。通过大语言模型(LLM)协调3D轨迹与摄像机运动,并将这些编排轨迹作为视频生成的控制信号,确保严格的物理逻辑与外观稳定性,成功在物体长时间离屏后重入场景时仍保持其精确的视觉身份。实验表明,该方法能以前所未有的可控性合成复杂且持续的事件,实现持久动态对象记忆。项目页面:https://worlddirector.github.io/

原文摘要 · Abstract (English)

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/

世界模型可控生成动态记忆3D轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。