用多视角隐空间扩散模型统一补全随意拍摄的3D场景缺失区域
Fillerbuster: Unified Generative Scene Completion Model for Casual Captures
- 基于多视图潜在扩散变换器,联合生成未知视角并恢复相机位姿
- 在两个数据集上实现部分捕获场景的完整重建,支持未标定输入
- 开源框架可集成至Nerfstudio或Gsplat,适配多模态扩展
我们提出Fillerbuster,一种统一的生成式场景补全模型,采用多视角隐空间扩散变换器完成3D场景中未知区域的重建。随意拍摄常导致数据稀疏,物体遮挡后方或上方内容缺失。现有方法多聚焦于稀疏视图先验下优化已知像素质量,或仅从一两张照片补全物体侧面,难以应对数百帧输入且需补全观测范围外区域的真实场景。我们的方案训练一个生成模型,可在大上下文输入下同时生成未知目标视图并恢复相机位姿(当参数未知时)。实验表明,该模型在两个现有数据集上成功完成部分捕获场景的补全,并引入未标定场景补全任务,实现位姿与新内容的联合预测。我们开源了该框架,可集成至Nerfstudio或Gsplat等主流重建平台。本工作提供灵活统一的图像与位姿联合修复框架,未来可扩展至深度等更多模态。
原文摘要 · Abstract (English)
We present Fillerbuster, a unified model that completes unknown regions of a 3D scene with a multi-view latent diffusion transformer. Casual captures are often sparse and miss surrounding content behind objects or above the scene. Existing methods are not suitable for this challenge as they focus on making known pixels look good with sparse-view priors, or on creating missing sides of objects from just one or two photos. In reality, we often have hundreds of input frames and want to complete areas that are missing and unobserved from the input frames. Our solution is to train a generative model that can consume a large context of input frames while generating unknown target views and recovering image poses when camera parameters are unknown. We show results where we complete partial captures on two existing datasets. We also present an uncalibrated scene completion task where our unified model predicts both poses and creates new content. We open-source our framework for integration into popular reconstruction platforms like Nerfstudio or Gsplat. We present a flexible, unified inpainting framework to predict many images and poses together, where all inputs are jointly inpainted, and it could be extended to predict more modalities such as depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。