arXiv:2603.11911cs.CV2026-03被引 6

实时生成高一致性三维场景帧,支持消费级显卡交互式探索。

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

  • 独立生成每帧,避免序列延迟,实现低时延实时推理。
  • 通过3D锚点与隐式记忆保持多视角几何一致性,细节保留好。
  • 三阶段蒸馏训练,可将图像扩散模型转为可控实时生成器。

我们提出InSpatio-WorldFM,一个开源的实时帧生成模型,用于空间智能。不同于依赖序列帧生成且因窗口处理导致显著延迟的视频类世界模型,InSpatio-WorldFM采用帧级生成范式,独立生成每帧,实现低延迟实时空间推断。通过显式3D锚点和隐式空间记忆强制多视角空间一致性,模型在视点变化下保持全局场景几何结构的同时保留精细视觉细节。我们进一步设计了一种渐进式三阶段训练流程,将预训练图像扩散模型逐步转化为可控帧模型,并最终变为实时生成器,仅需少量步骤蒸馏即可完成。实验表明,InSpatio-WorldFM在保证强多视角一致性的同时,可在消费级GPU上实现交互式探索,为实时世界模拟提供了高效替代方案。

原文摘要 · Abstract (English)

We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur substantial latency due to window-level processing, InSpatio-WorldFM adopts a frame-based paradigm that generates each frame independently, enabling low-latency real-time spatial inference. By enforcing multi-view spatial consistency through explicit 3D anchors and implicit spatial memory, the model preserves global scene geometry while maintaining fine-grained visual details across viewpoint changes. We further introduce a progressive three-stage training pipeline that transforms a pretrained image diffusion model into a controllable frame model and finally into a real-time generator through few-step distillation. Experimental results show that InSpatio-WorldFM achieves strong multi-view consistency while supporting interactive exploration on consumer-grade GPUs, providing an efficient alternative to traditional video-based world models for real-time world simulation.

生成模型实时推理空间建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。