arXiv:2504.02764cs.CVcs.AI2025-04CVPR被引 14

用视频扩散模型从单图生成连贯3D场景,解决画面不一致问题。

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

论文配图:Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
图 1 · 摘自论文原文
  • 引入动量机制,分层增强视频细节与场景一致性。
  • 在未知区域恢复上表现优异,生成视频长度无限制。
  • 适合需要高质量3D场景重建的视觉生成任务。

本文提出 Scene Splatter,一种基于动量的视频扩散方法,用于从单张图像生成通用3D场景。现有方法使用视频生成模型合成新视角时,受限于视频长度且易出现场景不一致,导致重建中产生伪影和扭曲。为解决此问题,我们从原始特征构造噪声样本作为动量,以提升视频细节并维持场景一致性。然而,对于感知范围涵盖已知与未知区域的潜在特征,这种潜层动量会限制视频扩散在未知区域的生成能力。因此,我们进一步将上述一致视频作为像素级动量,作用于无动量直接生成的视频,以更好恢复未见区域。该级联动量机制使视频扩散模型能生成高保真且一致的新视角。我们还通过增强帧微调全局高斯表示,并用于下一阶段的动量更新。如此可迭代重建3D场景,突破视频长度限制。大量实验表明,本方法在高保真度与一致性场景生成上具备优异泛化能力和性能。

原文摘要 · Abstract (English)

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suffer from limited video length and scene inconsistency, leading to artifacts and distortions during further reconstruction. To address this issue, we construct noisy samples from original features as momentum to enhance video details and maintain scene consistency. However, for latent features with the perception field that spans both known and unknown regions, such latent-level momentum restricts the generative ability of video diffusion in unknown regions. Therefore, we further introduce the aforementioned consistent video as a pixel-level momentum to a directly generated video without momentum for better recovery of unseen regions. Our cascaded momentum enables video diffusion models to generate both high-fidelity and consistent novel views. We further finetune the global Gaussian representations with enhanced frames and render new frames for momentum update in the next step. In this manner, we can iteratively recover a 3D scene, avoiding the limitation of video length. Extensive experiments demonstrate the generalization capability and superior performance of our method in high-fidelity and consistent scene generation.

3D生成视频扩散单图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。