arXiv:2504.07146eess.IV2025-04CVPR被引 1

用时空样条解耦视频的运动、外观和遮挡,实现更自然的视频编辑。

VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing

  • 用空间与颜色样条场分离运动与外观信息。
  • 在DAVIS、CoDeF等数据集上实现长期时序一致性。
  • 适合需要精细控制的视频编辑场景,如人脸视频处理。

我们提出一种隐式视频表示方法——时空样条视频(VideoSPatS),用于从单目视频中解耦遮挡、外观和运动。与以往将时间和坐标映射为形变与原始颜色的方法不同,VideoSPatS将输入坐标映射为空间样条形变场 $D_s$ 和颜色样条形变场 $D_c$,从而在视频中分离运动与外观。基于样条的参数化方式自然生成时序一致的光流,并保证长期时序一致性,这对逼真的视频编辑至关重要。通过多分支预测,模型还能实现潜在视频与选定遮挡物之间的层分离。解耦遮挡、外观和运动后,方法在多样化视频(包括存在复杂遮挡、阴影和高光的开源网络说话头视频)上实现了更优的时空建模与编辑效果,同时保持合适的编辑基底空间。我们在DAVIS、CoDeF数据集以及自建的说话头视频数据集上展示了通用视频建模结果。大量消融实验表明,$D_s$ 与 $D_c$ 在神经样条下的结合可克服运动与外观的模糊性,为更先进的视频编辑模型铺平道路。

原文摘要 · Abstract (English)

We present an implicit video representation for occlusions, appearance, and motion disentanglement from monocular videos, which we call Video SPatiotemporal Splines (VideoSPatS). Unlike previous methods that map time and coordinates to deformation and canonical colors, our VideoSPatS maps input coordinates into Spatial and Color Spline deformation fields $D_s$ and $D_c$, which disentangle motion and appearance in videos. With spline-based parametrization, our method naturally generates temporally consistent flow and guarantees long-term temporal consistency, which is crucial for convincing video editing. Using multiple prediction branches, our VideoSPatS model also performs layer separation between the latent video and the selected occluder. By disentangling occlusions, appearance, and motion, our method enables better spatiotemporal modeling and editing of diverse videos, including in-the-wild talking head videos with challenging occlusions, shadows, and specularities while maintaining an appropriate canonical space for editing. We also present general video modeling results on the DAVIS and CoDeF datasets, as well as our own talking head video dataset collected from open-source web videos. Extensive ablations show the combination of $D_s$ and $D_c$ under neural splines can overcome motion and appearance ambiguities, paving the way for more advanced video editing models.

视频编辑时空建模样条表示解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。