将渲染图直接注入噪声流,提升视频重拍摄的轨迹控制精度
Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting

- 把渲染结果融入初始噪声流,生成从带几何信息的流形出发
- 旋转误差降低25%-27%,平移误差最多降32%,优于最强基线
- 适合追求高轨迹精度和动态一致性的视频重拍摄应用
视频重拍摄通过用户指定的相机轨迹重新渲染单目动态场景视频。现有方法显式提供目标几何:每帧深度将源视频升维为4D点云,并沿轨迹光栅化生成点云渲染。由于渲染和源视频均作为视觉条件输入网络,二者在去噪步骤中竞争,导致模型产生信任困境——该信谁?这会损害分布外数据的轨迹控制或视觉质量。本文认为,已与目标视图像素对齐的渲染无需作为显式条件输入。提出MANIFOLD4D:将渲染直接注入流匹配的初始噪声中,使生成不再从标准高斯噪声开始,而是从携带几何信息的新噪声流形出发,仅以源视频作为视觉条件。渲染仅使用一次,模型无需学习如何解读它;后续去噪阶段可专注源视频。在DAVIS-Traj和Vista4D数据集上,MANIFOLD4D在所有指标上达到最优,旋转误差降低25%和27%,平移误差最高下降32%,视频保真度相当,真实世界新视角的光照一致性更优。用户研究表明,本方法在轨迹跟随和动态一致性上优势明显,且当偏航幅度超过训练范围时差距更大;即使渲染被故意破坏,模型仍能正确恢复源视频中的动态运动,验证了几何先验引导生成而不覆盖生成过程。
原文摘要 · Abstract (English)
Video re-shooting re-renders a monocular video of a dynamic scene along a user-specified camera trajectory, and the dominant recipe supplies the target geometry explicitly: per-frame depth lifts the source video into a 4D point cloud, which is rasterized along the trajectory into a point cloud render. Because the render and the source video are both handed to the network as visual conditions, they compete at every denoising step, leaving the model with a trust dilemma --- how much of the render to believe --- which can degrade trajectory control or visual quality on data outside the training distribution. We argue that a render already pixel-aligned with the target view does not need to be supplied as an explicit conditioning stream at all. We propose MANIFOLD4D, which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a new noise manifold carrying geometric information, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our DAVIS-Traj benchmark and on the Vista4D evaluation set, MANIFOLD4D attains the best camera-control accuracy on every metric, lowering rotation error by 25% and 27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and leading on real-world novel-view photometric quality. In a user study, our method achieves clear advantages in trajectory following and dynamic consistency. The gap widens as the yaw amplitude grows past the training range, and the model still recovers correct dynamic motion from the source video when the render is deliberately corrupted, confirming that the geometric prior guides generation without overriding it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。