用文本和参考图生成一致的360度3D场景,解决纹理差、结构不连贯问题。
PanoDreamer: Consistent Text to 360-Degree Scene Generation
- 结合大语言模型与图像扭曲-精修流程,分步生成全景图并构建初始点云。
- 通过多视角新增图像迭代优化点云,提升几何一致性与细节质量。
- 适合虚拟现实、游戏等需高保真360度场景生成的应用场景。
从文本描述、参考图像或两者结合自动生成完整3D场景,在虚拟现实与游戏等领域具有重要应用。然而,现有方法常出现纹理质量低、3D结构不一致的问题,尤其在显著超出参考图像视域范围时更为明显。为此,我们提出PanoDreamer,一种支持灵活文本与图像控制的一致性3D场景生成框架。该方法首先利用大语言模型与扭曲-精修流水线生成初始图像集,并拼接为360度全景图;再将其提升至3D形成初始点云。随后,基于该点云生成多个新视角图像,进一步扩展并精修点云。最终,使用3D高斯溅射(3D Gaussian Splatting)构建可从任意视角渲染的高质量3D场景。实验表明,PanoDreamer能有效生成高质量、几何一致的3D场景。
原文摘要 · Abstract (English)
Automatically generating a complete 3D scene from a text description, a reference image, or both has significant applications in fields like virtual reality and gaming. However, current methods often generate low-quality textures and inconsistent 3D structures. This is especially true when extrapolating significantly beyond the field of view of the reference image. To address these challenges, we propose PanoDreamer, a novel framework for consistent, 3D scene generation with flexible text and image control. Our approach employs a large language model and a warp-refine pipeline, first generating an initial set of images and then compositing them into a 360-degree panorama. This panorama is then lifted into 3D to form an initial point cloud. We then use several approaches to generate additional images, from different viewpoints, that are consistent with the initial point cloud and expand/refine the initial point cloud. Given the resulting set of images, we utilize 3D Gaussian Splatting to create the final 3D scene, which can then be rendered from different viewpoints. Experiments demonstrate the effectiveness of PanoDreamer in generating high-quality, geometrically consistent 3D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。