用分步扩展生成高保真全景3D场景,解决沉浸式建模中清晰度与可探索性的矛盾。
Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
- 分步扩展全景图,结合多视角扩散模型保持视觉一致性
- 在新数据集上训练,实现最高水平的图像保真度和结构一致性
- 适合做AR/VR内容生成、虚拟世界建模的研究者和开发者
从文本生成沉浸式3D场景正快速发展,得益于新型视频生成模型与前馈3D重建技术,在AR/VR和世界建模中有巨大潜力。尽管全景图已证明对场景初始化有效,现有方法在视觉保真度与可探索性之间存在权衡:自回归扩展易出现上下文漂移,而全景视频生成受限于低分辨率。我们提出Stepper,一种统一的文本驱动沉浸式3D场景生成框架,通过分步全景场景扩展规避上述限制。Stepper采用新颖的多视角360°扩散模型,实现一致且高分辨率的扩展,并结合几何重建流水线确保几何一致性。在新构建的大规模多视角全景数据集上训练,Stepper在保真度与结构一致性上均优于现有方法,树立了沉浸式场景生成的新标准。
原文摘要 · Abstract (English)
The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and world modeling. While panoramic images have proven effective for scene initialization, existing approaches suffer from a trade-off between visual fidelity and explorability: autoregressive expansion suffers from context drift, while panoramic video generation is limited to low resolution. We present Stepper, a unified framework for text-driven immersive 3D scene synthesis that circumvents these limitations via stepwise panoramic scene expansion. Stepper leverages a novel multi-view 360° diffusion model that enables consistent, high-resolution expansion, coupled with a geometry reconstruction pipeline that enforces geometric coherence. Trained on a new large-scale, multi-view panorama dataset, Stepper achieves state-of-the-art fidelity and structural consistency, outperforming prior approaches, thereby setting a new standard for immersive scene generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。