用单视频实时生成可交互的4D动态世界,空间一致且响应精准。
INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling
- 基于时空自回归架构,融合隐式缓存与显式约束实现场景持续演化。
- 在WorldScore-Dynamic上超越现有实时方法,导航一致性与交互精度领先。
- 适合需要高保真动态环境重建的虚拟现实、自动驾驶等应用。
构建具备空间一致性与实时交互能力的世界模型仍是计算机视觉中的核心挑战。当前视频生成方法常因缺乏空间持久性与视觉真实感,难以支持复杂环境中的无缝导航。为此,我们提出INSPATIO-WORLD,一种能从单个参考视频恢复并生成高保真、可交互动态场景的实时框架。其核心是时空自回归(STAR)架构,由两个紧密耦合组件构成:隐式时空缓存(Implicit Spatiotemporal Cache)将参考与历史观测聚合为潜在世界表示,确保长时程导航中的全局一致性;显式空间约束模块(Explicit Spatial Constraint Module)强化几何结构,并将用户操作转化为精确且物理合理的相机轨迹。此外,引入联合分布匹配蒸馏(JDMD),利用真实世界数据分布作为正则化引导,有效缓解过度依赖合成数据导致的保真度下降。大量实验表明,INSPATIO-WORLD在空间一致性与交互精度上显著优于现有最先进模型,在WorldScore-Dynamic基准上位居实时交互方法首位,为从单目视频重建4D环境提供了实用流程。
原文摘要 · Abstract (English)
Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insufficient visual realism, making it difficult to support seamless navigation in complex environments. To address these challenges, we propose INSPATIO-WORLD, a novel real-time framework capable of recovering and generating high-fidelity, dynamic interactive scenes from a single reference video. At the core of our approach is a Spatiotemporal Autoregressive (STAR) architecture, which enables consistent and controllable scene evolution through two tightly coupled components: Implicit Spatiotemporal Cache aggregates reference and historical observations into a latent world representation, ensuring global consistency during long-horizon navigation; Explicit Spatial Constraint Module enforces geometric structure and translates user interactions into precise and physically plausible camera trajectories. Furthermore, we introduce Joint Distribution Matching Distillation (JDMD). By using real-world data distributions as a regularizing guide, JDMD effectively overcomes the fidelity degradation typically caused by over-reliance on synthetic data. Extensive experiments demonstrate that INSPATIO-WORLD significantly outperforms existing state-of-the-art (SOTA) models in spatial consistency and interaction precision, ranking first among real-time interactive methods on the WorldScore-Dynamic benchmark, and establishing a practical pipeline for navigating 4D environments reconstructed from monocular videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。