用3D结构重建和渲染实现稳定视频,减少抖动且不裁剪画面。
GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and Rendering
- 基于3D相机位姿,通过局部重建与渲染保持时间一致性。
- 在多个数据集上优于或媲美现有2D/2.5D方法,几何一致性更优。
- 适合需要高质量视频稳定的应用,如无人机、手机拍摄。
视频稳定对视频处理至关重要,能消除抖动同时保留用户运动意图。现有方法因领域限制,常出现几何失真、过度裁剪、泛化能力差等问题。为此,我们提出GaVS,一种新型3D grounded方法,将视频稳定重构为时序一致的“局部重建与渲染”范式。给定3D相机位姿,我们增强重建模型以预测高斯点云(Gaussian Splatting)基础单元,并在测试时通过多视角动态感知光度监督与跨帧正则化进行微调,生成时序一致的局部重建。随后利用重建结果渲染每帧稳定视频。引入场景外推模块避免帧裁剪。我们在一个重新构建的数据集上评估,该数据集包含带有3D信息的多样化相机运动与场景动态样本。定量结果显示,本方法在传统任务指标及新提出的几何一致性指标上均优于或媲美当前最先进的2D与2.5D方法。定性分析表明,相比其他方法,本方法效果显著更优,用户研究验证了这一点。
原文摘要 · Abstract (English)
Video stabilization is pivotal for video processing, as it removes unwanted shakiness while preserving the original user motion intent. Existing approaches, depending on the domain they operate, suffer from several issues (e.g. geometric distortions, excessive cropping, poor generalization) that degrade the user experience. To address these issues, we introduce \textbf{GaVS}, a novel 3D-grounded approach that reformulates video stabilization as a temporally-consistent `local reconstruction and rendering' paradigm. Given 3D camera poses, we augment a reconstruction model to predict Gaussian Splatting primitives, and finetune it at test-time, with multi-view dynamics-aware photometric supervision and cross-frame regularization, to produce temporally-consistent local reconstructions. The model are then used to render each stabilized frame. We utilize a scene extrapolation module to avoid frame cropping. Our method is evaluated on a repurposed dataset, instilled with 3D-grounded information, covering samples with diverse camera motions and scene dynamics. Quantitatively, our method is competitive with or superior to state-of-the-art 2D and 2.5D approaches in terms of conventional task metrics and new geometry consistency. Qualitatively, our method produces noticeably better results compared to alternatives, validated by the user study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。