arXiv:2506.18601cs.GRcs.AI2025-06中稿 · CVPR被引 2

用生成模型修复单目视频的4D重建误差,实现更沉浸的动态场景还原。

BulletGen: Improving 4D Reconstruction with Bullet-Time Generation

  • 利用扩散模型生成子弹时间帧,指导4D高斯模型优化
  • 在新视角合成与2D/3D追踪任务上达到当前最优性能
  • 融合生成内容与真实场景,有效补全遮挡区域

将随意拍摄的单目视频转化为全沉浸式动态体验是一项高度病态的任务,面临诸多挑战,如重建不可见区域、解决单目深度估计的歧义性。本文提出BulletGen,一种利用生成模型校正基于高斯的动态场景表示中的误差并补全缺失信息的方法。通过将基于扩散的视频生成模型输出与某一固定“子弹时间”步骤下的4D重建对齐,生成帧被用于监督4D高斯模型的优化。该方法无缝融合生成内容与静态及动态场景成分,在新视角合成以及2D/3D追踪任务上均取得当前最优结果。

原文摘要 · Abstract (English)

Transforming casually captured, monocular videos into fully immersive dynamic experiences is a highly ill-posed task, and comes with significant challenges, e.g., reconstructing unseen regions, and dealing with the ambiguity in monocular depth estimation. In this work we introduce BulletGen, an approach that takes advantage of generative models to correct errors and complete missing information in a Gaussian-based dynamic scene representation. This is done by aligning the output of a diffusion-based video generation model with the 4D reconstruction at a single frozen "bullet-time" step. The generated frames are then used to supervise the optimization of the 4D Gaussian model. Our method seamlessly blends generative content with both static and dynamic scene components, achieving state-of-the-art results on both novel-view synthesis, and 2D/3D tracking tasks.

4D重建生成模型单目视频高斯表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。