用4D点云重拍视频,实现精准视角切换与动态保持。
Vista4D: Video Reshooting with 4D Point Clouds

- 以4D点云为基准,融合静态分割与动态重建,统一场景表征
- 在多类视频与轨迹上优于主流方法,视觉质量与相机控制更优
- 适用于真实场景扩展与4D重构,适合影视与虚拟拍摄应用
我们提出Vista4D,一种鲁棒且灵活的视频重拍框架,将输入视频与目标摄像机统一到4D点云空间中。给定一段输入视频,该方法可从新视角和路径重新合成具有相同动态特性的场景。现有方法常因真实动态视频的深度估计误差导致问题,且难以保持内容外观或对复杂轨迹进行精确相机控制。Vista4D通过静态像素分割与4D重建构建4D对齐的点云表示,显式保留可见内容并提供丰富相机信号,并利用多视角动态数据重建进行训练,增强对实际推理中点云伪影的鲁棒性。实验表明,相比当前最优基线,本方法在多种视频与摄像机路径下实现了更高的4D一致性、相机控制精度与视觉质量。此外,该方法可推广至真实场景扩展与4D场景重组等应用。
原文摘要 · Abstract (English)
We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing video reshooting methods often struggle with depth estimation artifacts of real-world dynamic videos, while also failing to preserve content appearance and failing to maintain precise camera control for challenging new trajectories. We build a 4D-grounded point cloud representation with static pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals, and we train with reconstructed multiview dynamic data for robustness against point cloud artifacts during real-world inference. Our results demonstrate improved 4D consistency, camera control, and visual quality compared to state-of-the-art baselines under a variety of videos and camera paths. Moreover, our method generalizes to real-world applications such as dynamic scene expansion and 4D scene recomposition. See our project page for results, code, and models: https://eyeline-labs.github.io/Vista4D
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。