无需微调即可实现视频扩散模型的精准摄像机控制。
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
- 采样阶段通过时序点云重构潜在空间,实现相机轨迹对齐。
- 生成视频在摄像机控制精度和画质上媲美训练方法。
- 适合希望快速部署相机控制且不改变预训练模型的用户。
精确的摄像机姿态控制对扩散模型视频生成至关重要。现有方法需在包含配对视频与摄像机姿态标注的额外数据集上进行微调,数据与计算成本高,且可能破坏预训练模型分布。我们提出Latent-Reframe,可在不微调的情况下实现预训练视频扩散模型的摄像机控制。该方法在采样阶段运行,保持高效并保留原始模型分布。通过时序点云重构视频帧潜在码以对齐输入摄像机轨迹,并结合潜在码修复与协调,优化模型潜在空间,确保高质量视频生成。实验表明,Latent-Reframe在摄像机控制精度和视频质量上达到或超过训练类方法,且无需额外数据微调。
原文摘要 · Abstract (English)
Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera pose annotations, which are both data-intensive and computationally costly, and can disrupt the pre-trained model distribution. We introduce Latent-Reframe, which enables camera control in a pre-trained video diffusion model without fine-tuning. Unlike existing methods, Latent-Reframe operates during the sampling stage, maintaining efficiency while preserving the original model distribution. Our approach reframes the latent code of video frames to align with the input camera trajectory through time-aware point clouds. Latent code inpainting and harmonization then refine the model latent space, ensuring high-quality video generation. Experimental results demonstrate that Latent-Reframe achieves comparable or superior camera control precision and video quality to training-based methods, without the need for fine-tuning on additional datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。