2B参数模型实现顶尖图文生成与编辑,训练策略远超大模型
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- 采用渐进式双任务强化学习提升指令理解和编辑一致性
- 20亿参数模型性能超越70亿和120亿参数的同类模型
- 支持统一多模态框架,适合需要高效生成编辑的应用场景
近期多模态模型在统一图像生成与编辑方面展现强大能力。然而,许多开源模型过度追求参数规模,忽视训练策略优化,制约效率与性能。本文提出基于SD3.5-Medium的20亿参数DiT模型UniPic2-SD3.5M-Kontext,实现顶尖图像生成与编辑能力,并无缝扩展为统一多模态框架。通过架构改进与高质量数据大规模预训练,实现文本到图像生成与编辑的联合能力。为增强指令遵循与编辑一致性,提出新型渐进式双任务强化学习(PDTR),实证表明不同任务的强化阶段相互促进,无负向干扰。经预训练与强化策略后,该模型在生成与编辑能力上优于参数量更大的BAGEL(7B)与Flux-Kontext(12B)。进一步,通过连接器将UniPic2-SD3.5M-Kontext与Qwen2.5-VL-7B联合训练,构建统一多模态模型UniPic2-Metaquery,集成理解、生成与编辑能力,在多样任务中达到顶级表现。该训练范式被正式命名为Skywork UniPic 2.0,验证其有效性和通用性。
原文摘要 · Abstract (English)
Recent advances in multimodal models have demonstrated impressive capabilities in unified image generation and editing. However, many prominent open-source models prioritize scaling model parameters over optimizing training strategies, limiting their efficiency and performance. In this work, we present UniPic2-SD3.5M-Kontext, a 2B-parameter DiT model based on SD3.5-Medium, which achieves state-of-the-art image generation and editing while extending seamlessly into a unified multimodal framework. Our approach begins with architectural modifications to SD3.5-Medium and large-scale pre-training on high-quality data, enabling joint text-to-image generation and editing capabilities. To enhance instruction following and editing consistency, we propose a novel Progressive Dual-Task Reinforcement strategy (PDTR), which effectively strengthens both tasks in a staged manner. We empirically validate that the reinforcement phases for different tasks are mutually beneficial and do not induce negative interference. After pre-training and reinforcement strategies, UniPic2-SD3.5M-Kontext demonstrates stronger image generation and editing capabilities than models with significantly larger generation parameters-including BAGEL (7B) and Flux-Kontext (12B). Furthermore, following the MetaQuery, we connect the UniPic2-SD3.5M-Kontext and Qwen2.5-VL-7B via a connector and perform joint training to launch a unified multimodal model UniPic2-Metaquery. UniPic2-Metaquery integrates understanding, generation, and editing, achieving top-tier performance across diverse tasks with a simple and scalable training paradigm. This consistently validates the effectiveness and generalizability of our proposed training paradigm, which we formalize as Skywork UniPic 2.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。