arXiv:2411.04924cs.CV2024-11NeurIPS被引 121

仅用5张稀疏图像实现360度全景视角合成,效果超越现有方法。

MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views

  • 将3D高斯溅射与视频扩散模型结合,直接在隐空间生成新视角。
  • 在DL3DV-10K数据集上实现360°视角合成,视觉质量领先。
  • 适用于少视角、全场景重建,适合虚拟现实与数字孪生应用。

我们提出MVSplat360,一种基于稀疏观测的前馈式360°新视角合成方法。该任务因输入视图重叠少、信息不足而高度病态,传统方法难以获得高质量结果。MVSplat360通过融合几何感知3D重建与时间一致性视频生成,将前馈3D高斯溅射(3DGS)模型重构的特征直接映射至预训练稳定视频扩散(SVD)模型的隐空间,作为姿态与视觉引导信号,驱动去噪过程生成逼真且三维一致的新视角。模型端到端可训练,支持仅需5张稀疏输入即可渲染任意视角。为评估性能,我们引入基于挑战性DL3DV-10K数据集的新基准,结果表明其在广角甚至360°新视角合成任务中显著优于现有最先进方法。在RealEstate10K数据集上的实验也验证了模型有效性。视频演示见项目主页:https://donydchen.github.io/mvsplat360。

原文摘要 · Abstract (English)

We introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for conventional methods to achieve high-quality results. Our MVSplat360 addresses this by effectively combining geometry-aware 3D reconstruction with temporally consistent video generation. Specifically, it refactors a feed-forward 3D Gaussian Splatting (3DGS) model to render features directly into the latent space of a pre-trained Stable Video Diffusion (SVD) model, where these features then act as pose and visual cues to guide the denoising process and produce photorealistic 3D-consistent views. Our model is end-to-end trainable and supports rendering arbitrary views with as few as 5 sparse input views. To evaluate MVSplat360's performance, we introduce a new benchmark using the challenging DL3DV-10K dataset, where MVSplat360 achieves superior visual quality compared to state-of-the-art methods on wide-sweeping or even 360° NVS tasks. Experiments on the existing benchmark RealEstate10K also confirm the effectiveness of our model. The video results are available on our project page: https://donydchen.github.io/mvsplat360.

新视角合成360度重建扩散模型3D高斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。