单张图像生成可交互漫游的全景视频世界
MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

- 先扩成全景图再构建3D高斯骨架,实现完整场景重建
- 支持8帧/秒实时漫游,单张4090显卡即可运行
- 适合虚拟拍摄、元宇宙场景快速构建
我们提出MoVerse,一种从单张窄视场图像实时生成可交互漫游场景的视频世界模型。该任务挑战在于输入仅覆盖环境一小部分,而交互漫游需完整的周围世界、持续几何结构、可控相机运动及时间一致的高保真观测。MoVerse通过分离世界构建与观察渲染来应对:首先利用拓扑感知扩散将输入扩展为重力对齐的360°全景图,填补视野缺失;接着通过全景几何感知残差预测将全景图提升为持久的3D高斯骨架,形成密集且可直接渲染的空间记忆;最后,基于高斯条件的视频渲染器将骨架在用户指定相机轨迹上的渲染结果转化为逼真的视频。为支持实时交互,我们训练双向扩散教师模型进行高质量条件渲染,并蒸馏为因果自回归学生模型以实现低延迟流式传输。该设计融合了显式3D表示的可控性与生成模型的感知质量。MoVerse可在单张NVIDIA RTX 4090 GPU上实现8~12帧/秒的实时场景漫游,展示了单图像世界创建与交互视频输出的实用路径。
原文摘要 · Abstract (English)
We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete surrounding world, persistent geometry, controllable camera motion, and temporally coherent high-fidelity observations. MoVerse addresses this problem by separating world construction from observation rendering. It first expands the input into a gravity-aligned 360$^\circ$ panorama with topology-aware diffusion, closing the missing field of view before 3D reasoning. It then lifts the panorama into a persistent 3D Gaussian scaffold using panoramic geometry-aware residual prediction, yielding a dense and directly renderable spatial memory. Finally, a Gaussian-conditioned video renderer translates scaffold renderings along user-specified camera trajectories into photorealistic video. To make this renderer practical for interaction, we train a bidirectional diffusion teacher for high-quality conditional rendering and distill it into a causal autoregressive student for bounded-latency streaming. This design combines the controllability and long-range consistency of explicit 3D representations with the perceptual quality of generative video models. MoVerse supports real-time scene roaming at 8~FPS on a single NVIDIA RTX~4090 GPU, demonstrating a practical path toward single-image world creation with interactive video output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。