arXiv:2411.05003cs.CVcs.AI2024-11CVPR被引 75

让用户视频支持新镜头运动,还能补全看不见的画面

ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning

论文配图:ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
图 1 · 摘自论文原文
  • 先用多视角扩散或深度点云生成带噪新视角视频
  • 再通过掩码视频微调恢复清晰连贯的重拍视频
  • 适合想重制视频素材的创作者或影视工作者

近期视频建模突破使得生成视频可控制摄像机轨迹。然而,这些方法无法直接应用于非模型生成的用户提供的视频。本文提出ReCapture,一种从单个用户视频生成具有新摄像机轨迹的新视频的方法。该方法可在不同角度和电影级运镜下重生成参考视频,保留原有场景动态。值得注意的是,该方法还能合理推断出原视频中不可见区域的内容。具体流程为:(1) 使用多视角扩散模型或基于深度的点云渲染生成带有噪声的新视角视频;(2) 利用提出的掩码视频微调技术将锚定视频重构为清晰且时间一致的重拍视频。

原文摘要 · Abstract (English)

Recently, breakthroughs in video modeling have allowed for controllable camera trajectories in generated videos. However, these methods cannot be directly applied to user-provided videos that are not generated by a video model. In this paper, we present ReCapture, a method for generating new videos with novel camera trajectories from a single user-provided video. Our method allows us to re-generate the reference video, with all its existing scene motion, from vastly different angles and with cinematic camera motion. Notably, using our method we can also plausibly hallucinate parts of the scene that were not observable in the reference video. Our method works by (1) generating a noisy anchor video with a new camera trajectory using multiview diffusion models or depth-based point cloud rendering and then (2) regenerating the anchor video into a clean and temporally consistent reangled video using our proposed masked video fine-tuning technique.

视频重拍扩散模型相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。