仅用一张照片生成逼真3D室内场景,解决视角一致性难题。
LPA3D: 3D Room-Level Scene Generation from In-the-Wild Images
- 构建局部对齐坐标系,从单图推断相机位姿
- 联合优化位姿预测与场景生成,提升生成质量
- 适合虚拟现实、机器人导航等需要真实3D场景的领域
从真实世界图像生成具有语义合理性和细节丰富外观的房间级室内场景,在VR、AR和机器人应用中至关重要。基于NeRF的生成方法展现出巨大潜力,但现有场景级生成方法仍需多视角、深度图或语义引导,无法仅依赖RGB图像。这是因为NeRF需要已知相机位姿,而单张图像难以准确估计复杂室内场景的全局位姿。为此,我们提出局部位姿对齐(LPA)框架——一种基于锚点的多局部坐标系统,以选定锚点为坐标原点。在此基础上,提出LPA-GAN,一种改进的基于NeRF的生成方法,通过特定修改估计LPA下的相机位姿先验,并联合优化位姿预测与场景生成过程。消融实验和与直接扩展的物体级生成方法对比表明本方法有效性。视觉对比显示,该方法在视图间一致性与语义合理性方面表现更优。
原文摘要 · Abstract (English)
Generating realistic, room-level indoor scenes with semantically plausible and detailed appearances from in-the-wild images is crucial for various applications in VR, AR, and robotics. The success of NeRF-based generative methods indicates a promising direction to address this challenge. However, unlike their success at the object level, existing scene-level generative methods require additional information, such as multiple views, depth images, or semantic guidance, rather than relying solely on RGB images. This is because NeRF-based methods necessitate prior knowledge of camera poses, which is challenging to approximate for indoor scenes due to the complexity of defining alignment and the difficulty of globally estimating poses from a single image, given the unseen parts behind the camera. To address this challenge, we redefine global poses within the framework of Local-Pose-Alignment (LPA) -- an anchor-based multi-local-coordinate system that uses a selected number of anchors as the roots of these coordinates. Building on this foundation, we introduce LPA-GAN, a novel NeRF-based generative approach that incorporates specific modifications to estimate the priors of camera poses under LPA. It also co-optimizes the pose predictor and scene generation processes. Our ablation study and comparisons with straightforward extensions of NeRF-based object generative methods demonstrate the effectiveness of our approach. Furthermore, visual comparisons with other techniques reveal that our method achieves superior view-to-view consistency and semantic normality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。