arXiv:2606.30045cs.CV2026-06

用隐式场景状态实现流畅相机控制探索,生成更连贯视频。

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

论文配图:Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
图 1 · 摘自论文原文
  • 将视频生成的滚动变量从帧隐向量改为固定长度的隐式场景状态
  • 在长序列上保持高一致性,推理效率优于现有方法
  • 无需预训练模型或3D重建器,适合自监督场景建模

交互式视频生成系统在相机控制的世界探索中,持续生成隐空间视频帧,将状态转移与高频观测合成耦合。我们提出「Walking in the Implicit」,一种以场景为中心的范式,将滚动变量从帧隐向量改为固定长度、可渲染的隐式状态,称为神经隐式场景(NIS)。该范式将交互生成分解为紧凑场景状态的随机转移和基于采样状态的姿态条件渲染。我们构建了NeuWorld:一个变压器变分自编码器(VAE)从稀疏姿态图像中学习局部锚定的NIS,一个扩散变压器则根据未来相机轨迹和几何感知的历史检索结果演化NIS。通过复用VAE编码器作为统一条件器,NeuWorld将相机、参考图像和历史线索映射到同一NIS模态,避免使用外部异构编码器。整个模型从零训练,仅使用公开姿态图像数据,无需预训练视频骨干或辅助3D重建器,在长时序上实现强一致性,并具备良好的推理效率。

原文摘要 · Abstract (English)

Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transition with high-frequency observation synthesis. We propose Walking in the Implicit, a scene-centric paradigm that changes the rollout variable from frame latents to a fixed-length, renderable implicit state, termed Neural Implicit Scene (NIS). This factorizes interactive generation into stochastic transition of a compact scene state and deterministic pose-conditioned rendering given the sampled state. We instantiate this paradigm as NeuWorld: a transformer VAE learns locally anchored NIS from sparse posed frames, and a diffusion transformer evolves NIS conditioned on future camera trajectories and geometry-aware retrieved history. By reusing the VAE encoder as a unified conditioner, NeuWorld maps camera, reference-image, and history cues into the same NIS modality, avoiding external heterogeneous encoders. Trained from scratch on public posed-view data without pretrained video backbones or auxiliary 3D reconstructors, NeuWorld achieves strong long-horizon consistency with favorable inference efficiency.

视频生成隐式表示扩散模型场景建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。