让视频生成同时支持精细控制与自由发挥,提升可控性与一致性。
Ctrl-VI: Controllable Video Synthesis via Variational Inference
- 用变分推断统一处理多种控制输入,融合多个生成模型
- 在固定控制下保持视频多样性,3D一致性优于以往方法
- 适合需要精确轨迹或镜头路径的视频创作场景
许多视频工作流程需要混合使用不同粒度的用户控制,从精确的4维物体轨迹和相机路径到粗略的文本提示,而现有视频生成模型通常仅针对固定输入格式训练。我们提出Ctrl-VI,一种可控制视频生成方法,能在指定元素上实现高可控性,同时对未明确指定的部分保持多样性。将任务建模为变分推断,以逼近组合分布,并利用多个视频生成骨干网络共同满足所有约束条件。为应对优化挑战,我们通过渐进式分布序列最小化逐步KL散度,并提出一种上下文条件因子分解技术,减少解空间中的模式,避免陷入局部最优。实验表明,相比先前方法,本方法生成的视频在可控性、多样性和3D一致性方面均有提升。
原文摘要 · Abstract (English)
Many video workflows benefit from a mixture of user controls with varying granularity, from exact 4D object trajectories and camera paths to coarse text prompts, while existing video generative models are typically trained for fixed input formats. We develop Ctrl-VI, a video synthesis method that addresses this need and generates samples with high controllability for specified elements while maintaining diversity for under-specified ones. We cast the task as variational inference to approximate a composed distribution, leveraging multiple video generation backbones to account for all task constraints collectively. To address the optimization challenge, we break down the problem into step-wise KL divergence minimization over an annealed sequence of distributions, and further propose a context-conditioned factorization technique that reduces modes in the solution space to circumvent local optima. Experiments suggest that our method produces samples with improved controllability, diversity, and 3D consistency compared to prior works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。