无需训练即可用任意条件图像生成可控运动视频
AnyI2V: Animating Any Conditional Image with Motion Control
- 通过用户定义的运动轨迹驱动任意条件图像动画化
- 支持网格、点云等非图像数据作为输入,扩展了生成模态
- 兼容风格迁移与编辑,适合需要精细控制的创作者
近年来,扩散模型在文本到视频(T2V)和图像到视频(I2V)生成方面取得显著进展。然而,如何有效融合动态运动信号与灵活的空间约束仍是挑战。现有T2V方法依赖文本提示,难以精确控制内容的空间布局;而现有I2V方法受限于真实图像输入,合成内容可编辑性差。尽管部分方法采用ControlNet引入图像条件,但通常缺乏显式运动控制且需高成本训练。为此,我们提出AnyI2V——一种无需训练的框架,可基于用户定义的运动轨迹动画化任意条件图像。该框架支持包括网格、点云在内的多种数据类型作为条件输入,超越ControlNet限制,实现更灵活的视频生成。同时支持混合条件输入,可通过LoRA与文本提示实现风格迁移与编辑。大量实验表明,AnyI2V在空间与运动控制视频生成中表现优越,提供了新视角。代码已公开于https://henghuiding.com/AnyI2V/。
原文摘要 · Abstract (English)
Recent advancements in video generation, particularly in diffusion models, have driven notable progress in text-to-video (T2V) and image-to-video (I2V) synthesis. However, challenges remain in effectively integrating dynamic motion signals and flexible spatial constraints. Existing T2V methods typically rely on text prompts, which inherently lack precise control over the spatial layout of generated content. In contrast, I2V methods are limited by their dependence on real images, which restricts the editability of the synthesized content. Although some methods incorporate ControlNet to introduce image-based conditioning, they often lack explicit motion control and require computationally expensive training. To address these limitations, we propose AnyI2V, a training-free framework that animates any conditional images with user-defined motion trajectories. AnyI2V supports a broader range of modalities as the conditional image, including data types such as meshes and point clouds that are not supported by ControlNet, enabling more flexible and versatile video generation. Additionally, it supports mixed conditional inputs and enables style transfer and editing via LoRA and text prompts. Extensive experiments demonstrate that the proposed AnyI2V achieves superior performance and provides a new perspective in spatial- and motion-controlled video generation. Code is available at https://henghuiding.com/AnyI2V/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。