arXiv:2411.10836cs.CV2024-11CVPR被引 37

让视频生成更可控、更连贯,支持多种控制方式。

AnimateAnything: Consistent and Controllable Animation for Video Generation

  • 用多尺度特征融合统一处理相机、文本和动作标注的控制信号。
  • 生成视频时引入光流作为运动先验,提升一致性与精确性。
  • 通过频域稳定模块减少大运动带来的闪烁问题,适合动画创作。

我们提出一种统一的可控视频生成方法 AnimateAnything,可在相机轨迹、文本提示和用户动作标注等多种条件下实现精准且一致的视频操作。具体地,我们设计了一个多尺度控制特征融合网络,构建不同条件下的通用运动表征,并显式将所有控制信息转换为逐帧光流。随后将光流作为运动先验,引导最终视频生成。此外,为缓解大尺度运动引起的闪烁问题,我们提出一种基于频率的稳定模块,通过保证视频在频域上的一致性来增强时间连贯性。实验表明,该方法优于当前最先进方法。更多细节及演示视频请见:https://yu-shaonian.github.io/Animate_Anything/

原文摘要 · Abstract (English)

We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multi-scale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the webpage: https://yu-shaonian.github.io/Animate_Anything/.

视频生成可控生成光流引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。