arXiv:2504.13074cs.CV2025-04被引 210

SkyReels-V2实现无限时长电影级视频生成,突破长度与质量瓶颈。

SkyReels-V2: Infinite-length Film Generative Model

  • 融合多模态大模型与分阶段训练,构建影视级视频结构表示
  • 通过强化学习与扩散强迫框架,生成长达数十秒的流畅动态视频
  • 适合影视创作、长视频生成研究者使用,支持专业级镜头语言理解

近期视频生成进展主要依赖于扩散模型和自回归框架,但仍面临提示遵循性、视觉质量、运动动态和时长之间的权衡:为提升时间视觉质量而牺牲运动动态,受限于5-10秒的视频时长以保证分辨率,且通用多模态大模型难以理解电影语法(如镜头构图、演员表情、摄像机运动),导致镜头感知生成不足。为此,我们提出SkyReels-V2——一个无限时长电影生成模型,融合多模态大模型(MLLM)、多阶段预训练、强化学习与扩散强迫框架。首先,设计综合视频结构表示,结合多模态大模型的通用描述与子专家模型的镜头语言;借助人工标注,训练统一视频字幕模型SkyCaptioner-V1以高效标注数据。其次,采用渐进式分辨率预训练,并进行四阶段后训练增强:初始概念均衡监督微调(SFT)提升基线质量;基于人工标注与合成失真数据的运动特定强化学习(RL)解决动态伪影;采用非递减噪声调度的扩散强迫框架,实现高效长视频合成;最终高质量SFT进一步优化视觉保真度。所有代码与模型已开源于https://github.com/SkyworkAI/SkyReels-V2。

原文摘要 · Abstract (English)

Recent advances in video generation have been driven by diffusion models and autoregressive frameworks, yet critical challenges persist in harmonizing prompt adherence, visual quality, motion dynamics, and duration: compromises in motion dynamics to enhance temporal visual quality, constrained video duration (5-10 seconds) to prioritize resolution, and inadequate shot-aware generation stemming from general-purpose MLLMs' inability to interpret cinematic grammar, such as shot composition, actor expressions, and camera motions. These intertwined limitations hinder realistic long-form synthesis and professional film-style generation. To address these limitations, we propose SkyReels-V2, an Infinite-length Film Generative Model, that synergizes Multi-modal Large Language Model (MLLM), Multi-stage Pretraining, Reinforcement Learning, and Diffusion Forcing Framework. Firstly, we design a comprehensive structural representation of video that combines the general descriptions by the Multi-modal LLM and the detailed shot language by sub-expert models. Aided with human annotation, we then train a unified Video Captioner, named SkyCaptioner-V1, to efficiently label the video data. Secondly, we establish progressive-resolution pretraining for the fundamental video generation, followed by a four-stage post-training enhancement: Initial concept-balanced Supervised Fine-Tuning (SFT) improves baseline quality; Motion-specific Reinforcement Learning (RL) training with human-annotated and synthetic distortion data addresses dynamic artifacts; Our diffusion forcing framework with non-decreasing noise schedules enables long-video synthesis in an efficient search space; Final high-quality SFT refines visual fidelity. All the code and models are available at https://github.com/SkyworkAI/SkyReels-V2.

视频生成扩散模型长视频电影生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。