一款可联合生成视频音频、支持多模态编辑的电影级视频模型
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
- 采用双流扩散架构,分别生成视频与同步音频,共享文本编码器
- 支持1080p/32FPS/15秒长视频生成,实现高质量影视级内容创作
- 统一处理生成、修复、编辑任务,适合影视制作与多模态创意人群
SkyReels V4 是一个统一的多模态视频基础模型,支持视频与音频的联合生成、修复与编辑。模型采用双流多模态扩散变换器(MMDiT)架构,一分支合成视频,另一分支生成时序对齐音频,共享基于多模态大语言模型(MLLM)的强大文本编码器。该模型可接受文本、图像、视频片段、掩码和音频参考等多种多模态指令。通过结合MLLM的多模态指令理解能力与视频分支的上下文学习,模型可在复杂条件下注入精细视觉引导;音频分支则利用音频参考指导声音生成。视频侧采用通道拼接形式,统一处理图像转视频、视频扩展、视频编辑等多种修复任务,并自然延伸至基于视觉参考的修复与编辑。支持最高1080p分辨率、32帧每秒、15秒时长,实现高保真、多镜头、电影级视频与同步音频生成。为提升长时高分辨率生成效率,引入联合低分辨率全序列生成与高分辨率关键帧生成策略,随后通过专用超分与插帧模型优化。据我们所知,SkyReels V4是首个同时支持多模态输入、联合音视频生成与统一处理生成、修复、编辑的视频基础模型,且在电影级分辨率与时长下保持强效率与高质量。
原文摘要 · Abstract (English)
SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branch synthesizes video and the other generates temporally aligned audio, while sharing a powerful text encoder based on the Multimodal Large Language Models (MLLM). SkyReels V4 accepts rich multi modal instructions, including text, images, video clips, masks, and audio references. By combining the MLLMs multi modal instruction following capability with in context learning in the video branch MMDiT, the model can inject fine grained visual guidance under complex conditioning, while the audio branch MMDiT simultaneously leverages audio references to guide sound generation. On the video side, we adopt a channel concatenation formulation that unifies a wide range of inpainting style tasks, such as image to video, video extension, and video editing under a single interface, and naturally extends to vision referenced inpainting and editing via multi modal prompts. SkyReels V4 supports up to 1080p resolution, 32 FPS, and 15 second duration, enabling high fidelity, multi shot, cinema level video generation with synchronized audio. To make such high resolution, long-duration generation computationally feasible, we introduce an efficiency strategy: Joint generation of low resolution full sequences and high-resolution keyframes, followed by dedicated super-resolution and frame interpolation models. To our knowledge, SkyReels V4 is the first video foundation model that simultaneously supports multi-modal input, joint video audio generation, and a unified treatment of generation, inpainting, and editing, while maintaining strong efficiency and quality at cinematic resolutions and durations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。