重建3D场景时,能精准生成未观测视角的逼真图像和深度图。
ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views

- 用3D高斯泼溅作中间表示,结合多视角扩散模型联合优化外观与几何。
- 在真实数据集上,新视角生成和深度估计均优于现有方法。
- 适合需要高精度几何一致性的3D重建应用,如自动驾驶、VR。
我们提出ReconSplat,一种前馈式3D场景重建模型,旨在解决未观测区域生成合理视图与几何一致性之间的长期权衡问题,实现几何对齐的新视角生成和清晰的深度估计。该方法基于3D高斯泼溅(3DGS)作为可微分的中间场景表示,并集成多视图潜在扩散模型(MV-LDM),该模型同时充当外观与场景几何的精修器和补全器。通过利用由前馈3DGS表示编码并投影至2D潜空间的变分3D潜特征,引导扩散过程以确保几何一致性。ReconSplat在真实世界基准数据集RealEstate10K和DL3DV-10K上生成了逼真的新视角图像和准确的深度图,在具有挑战性的外推设置中表现更优。尤其值得注意的是,该模型能够联合实现难以观测视角的外推,同时保持连贯且精确的场景几何结构。
原文摘要 · Abstract (English)
We introduce ReconSplat, a feed-forward model for 3D scene reconstruction that aims to address the longstanding trade-off between plausible view generation for unobserved regions and geometric consistency, providing both geometrically aligned novel views and sharp depth estimates. Our approach builds on 3D Gaussian splatting (3DGS) as an intermediate differentiable scene representation and integrates it with a multi-view latent diffusion model (MV-LDM) trained to act simultaneously as a refiner and an inpainter for appearance and scene geometry. We enforce geometric consistency by guiding the diffusion process with variational 3D latent features for appearance and geometry, encoded by the feed-forward 3DGS representation and rasterized to 2D latent space. ReconSplat produces both photorealistic novel views and accurate depth maps on real-world benchmarks, RealEstate10K and DL3DV-10K, outperforming existing methods in challenging extrapolation setups. Notably, ReconSplat allows the extrapolation of unseen and challenging viewpoints jointly with coherent and precise scene geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。