用视频扩散与一致性感知点云,实现更真实更快的单图3D生成。
GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting
- 通过时间连贯性逐步填充视频帧,提升多视角一致性。
- 引入3D点云一致性掩码,使生成结果在多视角下更一致。
- 适合需要高质量3D生成且关注速度与鲁棒性的研究者。
单图像3D场景重建因本质上的不适定性及输入约束有限而面临挑战。现有方法中,多视角生成模型虽基于3D一致数据集训练,但泛化能力差;3D场景补全框架依赖深度或3D平滑性,易产生跨视角不一致,影响质量与效率。为此,我们提出GaussVideoDreamer,融合图像、视频与3D生成优势:(1) 采用渐进式视频补全策略,利用时间连贯性提升多视角一致性并加速收敛;(2) 引入3D高斯点云一致性掩码,引导视频扩散模型生成符合3D一致性的多视角输出。整体流程包含几何感知初始化、不一致感知高斯点云渲染及渐进式视频补全。实验表明,本方法在LLaVA-IQA评分上提升32%,速度至少快2倍,且在多样化场景中表现稳健。
原文摘要 · Abstract (English)
Single-image 3D scene reconstruction presents significant challenges due to its inherently ill-posed nature and limited input constraints. Recent advances have explored two promising directions: multiview generative models that train on 3D consistent datasets but struggle with out-of-distribution generalization, and 3D scene inpainting and completion frameworks that suffer from cross-view inconsistency and suboptimal error handling, as they depend exclusively on depth data or 3D smoothness, which ultimately degrades output quality and computational performance. Building upon these approaches, we present GaussVideoDreamer, which advances generative multimedia approaches by bridging the gap between image, video, and 3D generation, integrating their strengths through two key innovations: (1) A progressive video inpainting strategy that harnesses temporal coherence for improved multiview consistency and faster convergence. (2) A 3D Gaussian Splatting consistency mask to guide the video diffusion with 3D consistent multiview evidence. Our pipeline combines three core components: a geometry-aware initialization protocol, Inconsistency-Aware Gaussian Splatting, and a progressive video inpainting strategy. Experimental results demonstrate that our approach achieves 32% higher LLaVA-IQA scores and at least 2x speedup compared to existing methods while maintaining robust performance across diverse scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。