arXiv:2502.02690cs.CVcs.AI2025-02被引 4

提出新模型实现视频生成的精细控制,理论可证明解耦效果

Controllable Video Generation with Provable Disentanglement

  • 分离静态与动态潜在变量,按最小变化原则建模
  • 在多个数据集上提升生成质量与控制精度,有效解耦视频概念
  • 适合需要精准调控视频内容的研究者和开发者

可控视频生成仍是重大挑战,尽管近年已能生成高质量、连贯的视频。现有方法多将视频视为整体,忽视细粒度时空关系,限制了控制精度与效率。本文提出可控视频生成对抗网络(CoVoGAN),通过解耦视频概念,实现对各概念的高效独立控制。基于最小变化原则,首先分离静态与动态潜在变量;再利用充分变化特性,实现动态潜在变量的逐项可识别性,从而支持解耦控制。我们提供严格的理论分析,证明该方法的可识别性。在此基础上,设计时间过渡模块以解耦潜在动态。为满足最小变化与充分变化原则,最小化动态潜变量维度并施加时间条件独立性。通过作为插件集成至GANs,在多个视频生成基准上进行大量定性与定量实验,结果表明本方法显著提升了不同真实场景下的生成质量与可控性。

原文摘要 · Abstract (English)

Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video generation treat the video as a whole, neglecting intricate fine-grained spatiotemporal relationships, which limits both control precision and efficiency. In this paper, we propose Controllable Video Generative Adversarial Networks (CoVoGAN) to disentangle the video concepts, thus facilitating efficient and independent control over individual concepts. Specifically, following the minimal change principle, we first disentangle static and dynamic latent variables. We then leverage the sufficient change property to achieve component-wise identifiability of dynamic latent variables, enabling disentangled control of video generation. To establish the theoretical foundation, we provide a rigorous analysis demonstrating the identifiability of our approach. Building on these theoretical insights, we design a Temporal Transition Module to disentangle latent dynamics. To enforce the minimal change principle and sufficient change property, we minimize the dimensionality of latent dynamic variables and impose temporal conditional independence. To validate our approach, we integrate this module as a plug-in for GANs. Extensive qualitative and quantitative experiments on various video generation benchmarks demonstrate that our method significantly improves generation quality and controllability across diverse real-world scenarios.

视频生成生成对抗网络可控生成解耦表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。