arXiv:2511.14993cs.CVcs.AI2025-11被引 9

Kandinsky 5.0推出多模型家族,支持高清图像与10秒视频生成。

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

  • 分三类模型:6B参数图像生成、2B轻量视频生成、19B高质量视频生成
  • 通过数据清洗与自监督微调,实现高分辨率图像与10秒视频生成
  • 开源代码与模型,适合研究者快速部署生成应用

本文介绍Kandinsky 5.0,一套用于高分辨率图像与10秒视频合成的先进基础模型家族。包含三类核心模型:Kandinsky 5.0 Image Lite(6B参数图像生成)、Kandinsky 5.0 Video Lite(2B参数文本到视频/图像到视频轻量模型)和Kandinsky 5.0 Video Pro(19B参数高质量视频生成)。报告详细阐述了多阶段训练流程中的数据收集、处理、过滤与聚类全生命周期,并引入自监督微调(SFT)与基于强化学习(RL)的后训练质量提升技术。通过新型架构、训练与推理优化,Kandinsky 5.0在人类评估中展现出高速生成与领先性能。作为大规模公开可获取的生成框架,其充分释放预训练潜力,适用于广泛生成任务。我们希望本报告及开源代码与训练权重能显著推动高质量生成模型的研究与普及。

原文摘要 · Abstract (English)

This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core line-up of models: Kandinsky 5.0 Image Lite - a line-up of 6B parameter image generation models, Kandinsky 5.0 Video Lite - a fast and lightweight 2B parameter text-to-video and image-to-video models, and Kandinsky 5.0 Video Pro - 19B parameter models that achieves superior video generation quality. We provide a comprehensive review of the data curation lifecycle - including collection, processing, filtering and clustering - for the multi-stage training pipeline that involves extensive pre-training and incorporates quality-enhancement techniques such as self-supervised fine-tuning (SFT) and reinforcement learning (RL)-based post-training. We also present novel architectural, training, and inference optimizations that enable Kandinsky 5.0 to achieve high generation speeds and state-of-the-art performance across various tasks, as demonstrated by human evaluation. As a large-scale, publicly available generative framework, Kandinsky 5.0 leverages the full potential of its pre-training and subsequent stages to be adapted for a wide range of generative applications. We hope that this report, together with the release of our open-source code and training checkpoints, will substantially advance the development and accessibility of high-quality generative models for the research community.

图像生成视频生成扩散模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。