让文字指令精准生成包含指定元素的视频,保持图像一致性。
SkyReels-A2: Compose Anything in Video Diffusion Transformers
- 用多元素联合嵌入模型实现图文对齐与全局连贯性
- 构建三元组数据集并优化推理速度与稳定性
- 首个开源商用级可控视频生成框架,适合影视创作
本文提出SkyReels-A2,一种可控制视频生成框架,能根据文本提示将任意视觉元素(如角色、物体、背景)组合成合成视频,同时严格保持每个元素与参考图像的一致性。该任务称为元素到视频(E2V),核心挑战在于保留各参考元素的保真度、确保场景构图连贯性以及生成自然结果。为此,我们设计了完整的数据流水线,构建用于训练的提示-参考-视频三元组;提出新型图像-文本联合嵌入模型,将多元素表征注入生成过程,在元素特异性一致性和全局连贯性之间取得平衡;优化推理流程以提升速度与输出稳定性。此外,我们引入精心构建的评估基准A2 Bench。实验表明,该框架能生成多样且高质量的视频,实现精确元素控制。SkyReels-A2是首个开源的商用级E2V生成模型,在性能上优于先进闭源商业模型,有望推动戏剧创作与虚拟电商等创意应用的发展。
原文摘要 · Abstract (English)
This paper presents SkyReels-A2, a controllable video generation framework capable of assembling arbitrary visual elements (e.g., characters, objects, backgrounds) into synthesized videos based on textual prompts while maintaining strict consistency with reference images for each element. We term this task elements-to-video (E2V), whose primary challenges lie in preserving the fidelity of each reference element, ensuring coherent composition of the scene, and achieving natural outputs. To address these, we first design a comprehensive data pipeline to construct prompt-reference-video triplets for model training. Next, we propose a novel image-text joint embedding model to inject multi-element representations into the generative process, balancing element-specific consistency with global coherence and text alignment. We also optimize the inference pipeline for both speed and output stability. Moreover, we introduce a carefully curated benchmark for systematic evaluation, i.e, A2 Bench. Experiments demonstrate that our framework can generate diverse, high-quality videos with precise element control. SkyReels-A2 is the first open-source commercial grade model for the generation of E2V, performing favorably against advanced closed-source commercial models. We anticipate SkyReels-A2 will advance creative applications such as drama and virtual e-commerce, pushing the boundaries of controllable video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。