arXiv:2506.12716cs.CV2025-06被引 1

用生成先验重建复杂动态场景,支持遮挡下的多物体视角合成。

Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

  • 分物体优化可变形高斯,结合扩散模型填补视角缺失区域。
  • 在真实视频上实现4D物体重建,点轨迹误差降低23%以上。
  • 适合做动态场景生成与跟踪的科研人员和工程师。

针对单目多物体视频中存在严重遮挡时生成动态4D场景的挑战,本文提出GenMOJO,一种将基于渲染的可变形3D高斯优化与生成先验相结合的新方法。现有模型在孤立物体的新视角合成上表现良好,但在复杂杂乱场景中泛化能力差。GenMOJO将场景分解为独立物体,对每个物体优化可微分的可变形高斯集合,从而利用以物体为中心的扩散模型推断新视角中未观测区域。通过联合高斯点云渲染,捕获物体间遮挡关系,并实现遮挡感知监督。为弥合以物体为中心的生成先验与视频全局帧中心坐标系之间的差距,GenMOJO采用可微变换,在统一框架内对齐生成与渲染约束。最终模型实现了时空连续的4D物体重建,从单目输入生成准确的2D与3D点轨迹。定量评估与感知人类实验表明,GenMOJO生成的新视角更真实,点轨迹精度优于现有方法。

原文摘要 · Abstract (English)

We tackle the challenge of generating dynamic 4D scenes from monocular, multi-object videos with heavy occlusions, and introduce GenMOJO, a novel approach that integrates rendering-based deformable 3D Gaussian optimization with generative priors for view synthesis. While existing models perform well on novel view synthesis for isolated objects, they struggle to generalize to complex, cluttered scenes. To address this, GenMOJO decomposes the scene into individual objects, optimizing a differentiable set of deformable Gaussians per object. This object-wise decomposition allows leveraging object-centric diffusion models to infer unobserved regions in novel viewpoints. It performs joint Gaussian splatting to render the full scene, capturing cross-object occlusions, and enabling occlusion-aware supervision. To bridge the gap between object-centric priors and the global frame-centric coordinate system of videos, GenMOJO uses differentiable transformations that align generative and rendering constraints within a unified framework. The resulting model generates 4D object reconstructions over space and time, and produces accurate 2D and 3D point tracks from monocular input. Quantitative evaluations and perceptual human studies confirm that GenMOJO generates more realistic novel views of scenes and produces more accurate point tracks compared to existing approaches.

4D重建高斯溅射生成模型视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。