用多视角扩散模型实现3D场景中物体的高质量、一致插入。
Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model
- 基于预训练视频扩散模型构建多视角生成框架,支持视图一致性
- 引入ControlNet控制模块,提升生成结果的可预测性与可控性
- 适用于需要精准3D物体插入的场景重建任务,如虚拟拍摄
在由高斯点云表示的3D内容中生成并插入新物体,是实现多样化场景重建的有效途径。现有方法依赖于SDS优化或单视角修复,常导致质量不佳。为此,我们提出一种针对高斯点云表示的3D内容对象插入新方法。该方法引入名为MVInpainter的多视角扩散模型,基于预训练的Stable Video Diffusion模型,实现视图一致的物体修复。在MVInpainter中,采用基于ControlNet的条件注入模块,实现受控且更可预测的多视角生成。生成多视角修复结果后,进一步提出一种掩码感知的3D重建技术,从稀疏修复视图中优化高斯点云重建。通过这些协同技术,本方法实现了多样化的生成结果,确保视图一致性和视觉和谐,并显著提升物体质量。大量实验表明,本方法优于现有方法。
原文摘要 · Abstract (English)
Generating and inserting new objects into 3D content is a compelling approach for achieving versatile scene recreation. Existing methods, which rely on SDS optimization or single-view inpainting, often struggle to produce high-quality results. To address this, we propose a novel method for object insertion in 3D content represented by Gaussian Splatting. Our approach introduces a multi-view diffusion model, dubbed MVInpainter, which is built upon a pre-trained stable video diffusion model to facilitate view-consistent object inpainting. Within MVInpainter, we incorporate a ControlNet-based conditional injection module to enable controlled and more predictable multi-view generation. After generating the multi-view inpainted results, we further propose a mask-aware 3D reconstruction technique to refine Gaussian Splatting reconstruction from these sparse inpainted views. By leveraging these fabricate techniques, our approach yields diverse results, ensures view-consistent and harmonious insertions, and produces better object quality. Extensive experiments demonstrate that our approach outperforms existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。