用几何引导实现高效高质的3D视频生成,避免画面闪烁。
Geometry-guided Online 3D Video Synthesis with Multi-View Temporal Consistency
- 用全局几何信息逐步优化深度图,提升一致性。
- 通过截断符号距离场累积深度,实现跨视角与时间一致。
- 在线运行,适合实时3D视频合成场景。
我们提出一种新型几何引导的在线3D视频视图合成方法,显著提升视图与时间一致性。传统方法虽能从密集多视角相机中生成高质量结果,但计算开销大;而选择性输入方法虽降低资源消耗,常导致多视角和时间不一致,如闪烁伪影。本方法利用全局几何信息指导基于图像的渲染流程,通过时间序列中的颜色差异掩码逐步精炼深度图,并在合成视图的图像空间中以截断符号距离场(truncated signed distance fields)累积深度表示。该深度表示具有视图与时间一致性,用于引导预训练融合网络,融合多个前向渲染输入视图图像。因此,网络输出在多视角和时间上均保持几何一致。本方法在保证高质量视频合成的同时,支持高效在线运行。
原文摘要 · Abstract (English)
We introduce a novel geometry-guided online video view synthesis method with enhanced view and temporal consistency. Traditional approaches achieve high-quality synthesis from dense multi-view camera setups but require significant computational resources. In contrast, selective-input methods reduce this cost but often compromise quality, leading to multi-view and temporal inconsistencies such as flickering artifacts. Our method addresses this challenge to deliver efficient, high-quality novel-view synthesis with view and temporal consistency. The key innovation of our approach lies in using global geometry to guide an image-based rendering pipeline. To accomplish this, we progressively refine depth maps using color difference masks across time. These depth maps are then accumulated through truncated signed distance fields in the synthesized view's image space. This depth representation is view and temporally consistent, and is used to guide a pre-trained blending network that fuses multiple forward-rendered input-view images. Thus, the network is encouraged to output geometrically consistent synthesis results across multiple views and time. Our approach achieves consistent, high-quality video synthesis, while running efficiently in an online manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。