用3D感知的隐空间实现实时高质量新视角合成
LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
- 基于预训练3D重建网络初始化编码器,融合3D先验知识
- 在Re10k数据集上达到31.4 PSNR,支持实时渲染
- 适用于真实场景,可与扩散模型结合生成新视角
近期工作表明,神经网络可在无需显式3D重建的情况下完成3D任务如新视角合成(NVS)。然而我们认为,强3D归纳偏置对这类网络设计仍具价值。本文提出LagerNVS,一种基于编码器-解码器结构的NVS神经网络,利用‘3D感知’的隐空间特征。编码器由使用显式3D监督预训练的3D重建网络初始化,搭配轻量级解码器,并通过光度损失端到端训练。LagerNVS在无相机已知条件下也实现最优确定性前馈新视角合成,在Re10k数据集上达31.4 PSNR,支持实时渲染,能泛化至真实场景数据,并可与扩散解码器结合用于生成式外推。
原文摘要 · Abstract (English)
Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we argue that strong 3D inductive biases are still helpful in the design of such networks. We show this point by introducing LagerNVS, an encoder-decoder neural network for NVS that builds on `3D-aware' latent features. The encoder is initialized from a 3D reconstruction network pre-trained using explicit 3D supervision. This is paired with a lightweight decoder, and trained end-to-end with photometric losses. LagerNVS achieves state-of-the-art deterministic feed-forward Novel View Synthesis (including 31.4 PSNR on Re10k), with and without known cameras, renders in real time, generalizes to in-the-wild data, and can be paired with a diffusion decoder for generative extrapolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。