仅用一张图生成可自由漫游的3D场景,支持视角变换与时空一致。
NavCrafter: Exploring 3D Scenes from a Single Image
- 基于视频扩散模型与几何引导扩展,逐步拓展场景覆盖范围。
- 在大幅视角变化下实现顶尖的新视角合成效果,3D重建精度显著提升。
- 适合需要低成本3D内容生成的研究者与创作者使用。
从单张图像生成灵活的3D场景在直接获取3D数据成本高或不现实时至关重要。我们提出NavCrafter,一种新框架,通过合成具有相机可控性与时空一致性的新视角视频序列,实现从单图探索3D场景。NavCrafter利用视频扩散模型捕捉丰富的3D先验,并采用几何感知扩展策略逐步扩展场景覆盖范围。为实现可控多视角合成,引入多阶段相机控制机制,通过双分支相机注入与注意力调制,以多样化轨迹条件扩散模型。进一步提出碰撞感知相机轨迹规划器,以及改进的3D高斯溅射(3DGS)管道,包含深度对齐监督、结构正则化与精炼。大量实验表明,NavCrafter在大视角偏移下实现最先进的新视角合成效果,显著提升3D重建保真度。
原文摘要 · Abstract (English)
Creating flexible 3D scenes from a single image is vital when direct 3D data acquisition is costly or impractical. We introduce NavCrafter, a novel framework that explores 3D scenes from a single image by synthesizing novel-view video sequences with camera controllability and temporal-spatial consistency. NavCrafter leverages video diffusion models to capture rich 3D priors and adopts a geometry-aware expansion strategy to progressively extend scene coverage. To enable controllable multi-view synthesis, we introduce a multi-stage camera control mechanism that conditions diffusion models with diverse trajectories via dual-branch camera injection and attention modulation. We further propose a collision-aware camera trajectory planner and an enhanced 3D Gaussian Splatting (3DGS) pipeline with depth-aligned supervision, structural regularization and refinement. Extensive experiments demonstrate that NavCrafter achieves state-of-the-art novel-view synthesis under large viewpoint shifts and substantially improves 3D reconstruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。