Kandinsky 3 能用文本生成高清图像,并支持多种创作任务。
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
- 基于潜在扩散架构,统一支持文本、图像生成与编辑
- 4步推理速度提升3倍,图像质量仍属开源前列
- 适合需要多场景生成的创作者与研究者
文本到图像(T2I)扩散模型广泛用于图像编辑、融合、修复等任务。同时,图像到视频(I2V)和文本到视频(T2V)模型也建立在T2I基础上。本文提出Kandinsky 3,一种基于潜在扩散的新一代T2I模型,在质量与逼真度上达到高水平。其核心优势在于架构简洁高效,可灵活扩展至多种生成任务:包括文本引导的修复/外扩、图像融合、文图融合、图像变体生成、I2V与T2V生成。我们还推出了蒸馏版模型,在4步逆过程推理下保持高质量,速度比基础模型快3倍。已部署用户友好的公共演示系统,所有功能可在线体验。同时开源了Kandinsky 3及扩展模型的源代码与检查点。人类评估显示,Kandinsky 3在开源生成系统中质量评分居于前列。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models are popular for introducing image manipulation methods, such as editing, image fusion, inpainting, etc. At the same time, image-to-video (I2V) and text-to-video (T2V) models are also built on top of T2I models. We present Kandinsky 3, a novel T2I model based on latent diffusion, achieving a high level of quality and photorealism. The key feature of the new architecture is the simplicity and efficiency of its adaptation for many types of generation tasks. We extend the base T2I model for various applications and create a multifunctional generation system that includes text-guided inpainting/outpainting, image fusion, text-image fusion, image variations generation, I2V and T2V generation. We also present a distilled version of the T2I model, evaluating inference in 4 steps of the reverse process without reducing image quality and 3 times faster than the base model. We deployed a user-friendly demo system in which all the features can be tested in the public domain. Additionally, we released the source code and checkpoints for the Kandinsky 3 and extended models. Human evaluations show that Kandinsky 3 demonstrates one of the highest quality scores among open source generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。