arXiv:2601.17323cs.CV2026-01被引 14

SkyReels-V3统一生成三类视频,支持图像、视频和音频条件输入。

SkyReels-V3 Technique Report

论文配图:SkyReels-V3 Technique Report
图 1 · 摘自论文原文
  • 基于扩散Transformer的统一框架,融合多模态上下文学习。
  • 在视觉质量与指令遵循上达顶尖水平,接近闭源系统性能。
  • 适合需要高质量可控视频生成的研究者与开发者使用。

视频生成是构建世界模型的核心,多模态上下文推理成为能力的关键测试标准。为此,我们提出SkyReels-V3,一个基于统一多模态上下文学习框架的条件视频生成模型,采用扩散Transformer架构。该模型在同一架构内支持三大生成范式:参考图像生成视频、视频续写及音频引导视频生成。(i)参考图像到视频模型旨在生成高保真视频,保持主体身份一致、时间连贯性和叙事一致性。通过跨帧配对、图像编辑与语义重写的数据处理流程,有效缓解复制粘贴伪影。训练中采用图像-视频混合策略与多分辨率联合优化,提升多样场景下的泛化性与鲁棒性。(ii)视频扩展模型结合时空一致性建模与大规模视频理解,实现无缝单次续写与具备专业运镜模式的智能多段切换。(iii)Talker avatar模型通过首尾帧插入训练与关键帧推理重构,实现分钟级音频条件视频生成,同步性与画质均经优化。大量评估表明,SkyReels-V3在视觉质量、指令遵循与特定指标上达到或接近当前最优水平,逼近领先闭源系统。代码已开源:https://github.com/SkyworkAI/SkyReels-V3。

原文摘要 · Abstract (English)

Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a unified multimodal in-context learning framework with diffusion Transformers. SkyReels-V3 model supports three core generative paradigms within a single architecture: reference images-to-video synthesis, video-to-video extension and audio-guided video generation. (i) reference images-to-video model is designed to produce high-fidelity videos with strong subject identity preservation, temporal coherence, and narrative consistency. To enhance reference adherence and compositional stability, we design a comprehensive data processing pipeline that leverages cross frame pairing, image editing, and semantic rewriting, effectively mitigating copy paste artifacts. During training, an image video hybrid strategy combined with multi-resolution joint optimization is employed to improve generalization and robustness across diverse scenarios. (ii) video extension model integrates spatio-temporal consistency modeling with large-scale video understanding, enabling both seamless single-shot continuation and intelligent multi-shot switching with professional cinematographic patterns. (iii) Talking avatar model supports minute-level audio-conditioned video generation by training first-and-last frame insertion patterns and reconstructing key-frame inference paradigms. On the basis of ensuring visual quality, synchronization of audio and videos has been optimized. Extensive evaluations demonstrate that SkyReels-V3 achieves state-of-the-art or near state-of-the-art performance on key metrics including visual quality, instruction following, and specific aspect metrics, approaching leading closed-source systems. Github: https://github.com/SkyworkAI/SkyReels-V3.

视频生成扩散模型多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。