用生成模型提升3D场景重建质量,实现图像到3D的端到端生成。
Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction
- 通过适配器将重建模型输出转为几何潜在表示,与视频扩散模型对齐。
- 在单图和多图条件下均达到当前最优3D场景生成效果。
- 适合需要高质量3D重建与生成的视觉算法研究者。
我们提出Gen3R,一种将基础重建模型与视频扩散模型的强先验相结合的方法,用于场景级3D生成。通过在VGGT重建模型的令牌上训练适配器,生成几何潜在表示,并对其施加正则化以对齐预训练视频扩散模型的外观潜在空间。通过联合生成解耦但对齐的潜在变量,Gen3R同时输出RGB视频及对应的3D几何信息,包括相机位姿、深度图和全局点云。实验表明,该方法在单图与多图条件下的3D场景生成任务中均达到当前最优性能。此外,该方法通过引入生成先验提升了重建鲁棒性,验证了重建与生成模型紧密耦合的互惠价值。
原文摘要 · Abstract (English)
We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an adapter on its tokens, which are regularized to align with the appearance latents of pre-trained video diffusion models. By jointly generating these disentangled yet aligned latents, Gen3R produces both RGB videos and corresponding 3D geometry, including camera poses, depth maps, and global point clouds. Experiments demonstrate that our approach achieves state-of-the-art results in single- and multi-image conditioned 3D scene generation. Additionally, our method can enhance the robustness of reconstruction by leveraging generative priors, demonstrating the mutual benefit of tightly coupling reconstruction and generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。