让视频相机轨迹编辑更连贯,支持长时一致性生成。
CamDirector: Towards Long-Term Coherent Video Trajectory Editing
- 用世界缓存融合全局信息,分动静区域处理画面
- 通过历史引导扩散模型实现长序列时序一致
- 新基准测试表现领先,参数更少
视频相机轨迹编辑旨在生成沿用户定义相机路径的新视频,同时保留场景内容并合理补全未见区域,将业余视频升级为专业风格。现有方法在精确控制和长程一致性方面存在困难,因受限于容量的嵌入或仅依赖单帧变形且缺乏显式跨帧聚合。为此,我们提出新框架:1)通过混合变形方案显式聚合整个源视频信息,静态区域逐步融合至世界缓存后渲染到目标相机位姿,动态区域直接变形,融合生成全局一致的粗略帧以指导精修;2)通过历史引导自回归扩散模型联合处理视频片段及其历史,世界缓存增量更新以强化已补全内容,实现长期时序一致性。最后,我们构建了新基准iPhone-PTZ,涵盖多样相机运动与大轨迹变化,以更少参数达到当前最佳性能。
原文摘要 · Abstract (English)
Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regions, upgrading amateur footage into professionally styled videos. Existing VTE methods struggle with precise camera control and long-range consistency because they either inject target poses through a limited-capacity embedding or rely on single-frame warping with only implicit cross-frame aggregation in video diffusion models. To address these issues, we introduce a new VTE framework that 1) explicitly aggregates information across the entire source video via a hybrid warping scheme. Specifically, static regions are progressively fused into a world cache then rendered to target camera poses, while dynamic regions are directly warped; their fusion yields globally consistent coarse frames that guide refinement. 2) processes video segments jointly with their history via a history-guided autoregressive diffusion model, while the world cache is incrementally updated to reinforce already inpainted content, enabling long-term temporal coherence. Finally, we present iPhone-PTZ, a new VTE benchmark with diverse camera motions and large trajectory variations, and achieve state-of-the-art performance with fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。