用视频修复提升单目视频4D重建质量,融合几何与生成先验。
Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting

- 将多视角生成转化为视频修复任务,利用单目深度先验进行视图变换
- 在未标定网络视频上训练修复模型,合成遮挡掩码模拟视图扭曲
- 迭代增强视图并设计鲁棒损失,有效缓解单目深度误差
从随意拍摄的单目视频中重建4D动态场景具有价值但极具挑战性,因每个时刻仅从单一视角观测。我们提出Vivid4D,一种通过增强观测视角来提升单目视频4D合成的新方法——从单目输入合成多视角视频。不同于仅依赖几何先验或仅使用生成先验而忽略几何的方法,我们同时整合两者。这将视图增强重新建模为视频修复任务,基于单目深度先验将已观测视图扭曲至新视角。为此,我们在无姿态网络视频上训练视频修复模型,使用合成遮挡掩码模拟视图扭曲造成的遮挡,确保缺失区域的空间与时间一致性重建。为进一步缓解单目深度先验的不准确性,我们引入迭代视图增强策略和鲁棒重建损失。实验表明,该方法显著提升了单目4D场景重建与补全效果。
原文摘要 · Abstract (English)
Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D monocular video synthesis by augmenting observation views - synthesizing multi-view videos from a monocular input. Unlike existing methods that either solely leverage geometric priors for supervision or use generative priors while overlooking geometry, we integrate both. This reformulates view augmentation as a video inpainting task, where observed views are warped into new viewpoints based on monocular depth priors. To achieve this, we train a video inpainting model on unposed web videos with synthetically generated masks that mimic warping occlusions, ensuring spatially and temporally consistent completion of missing regions. To further mitigate inaccuracies in monocular depth priors, we introduce an iterative view augmentation strategy and a robust reconstruction loss. Experiments demonstrate that our method effectively improves monocular 4D scene reconstruction and completion. See our project page: https://xdimlab.github.io/Vivid4D/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。