用3D八叉树记忆生成长距离连贯视频,画面稳定不穿模。
OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping

- 用动态稀疏八叉树构建可扩展的3D视觉记忆
- 支持长镜头拍摄,重访区域仍保持空间一致性
- 适合需要真实感场景漫游的视频生成任务
我们提出OctWorld,一种具备持久3D记忆的视频扩散框架,可生成可探索、世界一致且高保真的视觉场景。输入单张图像后,OctWorld沿用户指定的相机轨迹进行稳定自回归世界生成。聚焦长距离生成,即路径长、视角广,此时重访区域的空间一致性尤为难维持。为此,我们引入OctMap——一种可扩展、空间自适应的3D记忆,将生成的视觉观测及其对应深度图逐步融合为全局表示。OctMap在动态稀疏八叉树中采用TSDF融合,空间分辨率随图像证据自适应调整。该设计在不同场景尺度下保持几何与外观细节,同时内存开销低。实验表明,OctWorld能生成长距离空间一致的视频,在现有基准和挑战性长程生成设置上均优于先前方法。OctMap相比点云缓存和固定分辨率TSDF体积有明显优势。
原文摘要 · Abstract (English)
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: https://maxtirerror.github.io/octworldpage/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。