Wonder让视频世界模型可实时操控相机探索,支持长期连贯的交互式场景生成。
Wonder: Video World Model Done Better

- 用密集坐标场实现相机运动的视觉化控制,直接理解镜头移动
- 16帧/秒生成细节丰富的视频,长时滚动保持几何与动态一致
- 适合需要实时交互式视频生成的研究者和开发者
我们提出Wonder,一种通用的视频世界模型,支持实时、可相机控制的世界探索。给定一张图像或条件视频,Wonder 构建出可交互的可玩世界,用户可通过移动相机实时探索未知区域、回看已观察区域,且支持长期连续推演。该能力依赖于控制方法、记忆机制与训练策略的系统级协同设计。我们引入新型相机条件机制,采用密集坐标场生成空间对齐的运动与朝向提示,使模型将相机运动直接视为视觉证据。为支持快速精准的记忆检索,提出基于稀疏注意力的记忆机制,推理时仅关注少量相关上下文标记,不随实际上下文长度增长而变慢。此外,改进自强化式蒸馏流程,提升学生模型对控制信号的响应能力,同时保留教师模型的多样生成模式与长期记忆。上述组件共同使Wonder在16 FPS下合成多样化、微观尺度的视频,长期推演中保持几何、外观与动态的一致性。除图像到视频生成外,Wonder也天然支持视频条件生成,实现现有动态场景的实时重拍。
原文摘要 · Abstract (English)
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。