用3D代理生成可交互场景,解决图像转场景的失真与漂移问题。
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

- 分两阶段生成:先建3D结构代理,再细化视觉细节。
- 在115K场景数据集上表现优于现有方法,视角变化下仍保持稳定。
- 适合游戏、机器人和虚拟现实中的高保真场景生成任务。
可交互的图像到场景生成对游戏、机器人和虚拟现实应用至关重要。现有基于视频扩散模型(VDM)的方法通常依赖稀疏点云或2D全景图等不完整条件信号,导致随机幻觉、长期漂移及三维一致性不足。我们提出SpatialCrafter,一种新颖的两阶段框架,通过引入全局3D代理实现高保真图像到场景生成。具体地,将生成过程分解为全局代理生成与外观精细化两个阶段。代理生成阶段提出点锚定稀疏结构(PaSS)流模块,预测空间对齐且几何一致的3D代理;外观精细化阶段将VDM重构为生成式延迟精炼器,在代理定义的场景几何基础上合成高频逼真细节。为更好融合代理与预训练VDM,引入并行几何注入与代理感知扰动训练策略,提升对代理伪影的鲁棒性,同时不破坏预训练生成流形。此外,由于缺乏合适的数据集,我们构建了包含115,000个场景的新大规模数据集,据我们所知,这是首个用于图像到场景生成的混合数据集。在合成与真实世界数据集上的大量实验表明,SpatialCrafter超越现有最优方法,缓解长期漂移,并在快速相机运动和极端视角变化下保持鲁棒性和一致性。
原文摘要 · Abstract (English)
Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency. We present SpatialCrafter, a novel two-stage framework that addresses these issues by introducing a global 3D proxy for high-fidelity image-to-scene generation. Specifically, we decompose the generation process into global proxy generation and appearance refinement. For proxy generation, we propose a Point-anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy. For appearance refinement, we re-frame the VDM as a Generative Deferred Refiner which synthesizes high-frequency photorealistic details upon proxy-defined scene geometry. To better integrate the proxy with the pre-trained VDM, we introduce Parallel Geometry Injection and Proxy-Aware Corruption training strategies, which improve robustness to proxy artifacts without disrupting the pretrained generative manifold. Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation. Extensive experiments on both synthetic and real-world datasets show that SpatialCrafter outperforms state-of-the-art methods, mitigates long-term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes. Our project page: \href{https://fangchuan.github.io/SpatialCrafter/}{fangchuan.github.io/SpatialCrafter/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。