一个框架搞定视频生成与编辑,支持多种任务组合。
VACE: All-in-One Video Creation and Editing

- 统一用视频条件单元处理不同任务输入。
- 在多个子任务上达到专用模型水平性能。
- 适合需要灵活视频创作的开发者和研究者。
扩散变换器在生成高质量图像和视频方面展现出强大能力与可扩展性。进一步推动生成与编辑任务的统一已取得显著进展。然而,由于时空动态一致性带来的内在挑战,实现统一的视频合成仍具难度。我们提出VACE,可在一体化框架中完成视频创作与编辑任务,包括参考视频生成、视频到视频编辑以及掩码视频到视频编辑。具体地,通过将编辑、参考和掩码等任务输入组织为统一接口——视频条件单元(VCU),有效整合各类任务需求。同时,借助上下文适配器结构,以时序与空间维度的形式化表示注入不同任务概念,使模型能灵活应对任意视频合成任务。大量实验表明,VACE的统一模型在各项子任务上性能与专用模型相当,同时支持多样化任务组合,拓展应用范围。
原文摘要 · Abstract (English)
Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of image content creation. However, due to the intrinsic demands for consistency across both temporal and spatial dynamics, achieving a unified approach for video synthesis remains challenging. We introduce VACE, which enables users to perform Video tasks within an All-in-one framework for Creation and Editing. These tasks include reference-to-video generation, video-to-video editing, and masked video-to-video editing. Specifically, we effectively integrate the requirements of various tasks by organizing video task inputs, such as editing, reference, and masking, into a unified interface referred to as the Video Condition Unit (VCU). Furthermore, by utilizing a Context Adapter structure, we inject different task concepts into the model using formalized representations of temporal and spatial dimensions, allowing it to handle arbitrary video synthesis tasks flexibly. Extensive experiments demonstrate that the unified model of VACE achieves performance on par with task-specific models across various subtasks. Simultaneously, it enables diverse applications through versatile task combinations. Project page: https://ali-vilab.github.io/VACE-Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。