统一多图合成框架,8步生成高质量图像,速度提升12.5倍。
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
- 将多图合成视为序列建模问题,统一单图与多图编辑任务。
- 仅用70万样本训练,支持1~6张输入图,输出分辨率灵活。
- 引入轨迹映射与分布匹配,推理仅需8步,速度提升12.5倍。
Nano-Banana和Seedream 4.0的流行凸显了社区对多图合成任务的强烈兴趣。相比单图编辑,多图合成在一致性与质量上更具挑战,但现有模型未公开具体实现方法。我们通过统计分析发现,人-物交互(HOI)是用户最关注的类别。为此,我们系统性地构建并实现了面向HOI任务的先进多图合成方案。提出Skywork UniPic 3.0,一个统一的多模态框架,支持1~6张任意分辨率输入,输出分辨率在总像素预算1024x1024内可调。设计完整的数据收集、过滤与合成流程,仅用70万高质量样本即取得优异性能。引入新型训练范式,将多图合成转化为序列建模问题,实现条件生成的统一序列合成。为加速推理,引入轨迹映射与分布匹配,使模型仅需8步即可生成高保真样本,较标准采样提速12.5倍。在单图编辑基准上达到顶尖水平,在多图合成基准上超越Nano-Banana与Seedream 4.0,验证了数据管道与训练范式的有效性。代码、模型与数据集已公开。
原文摘要 · Abstract (English)
The recent surge in popularity of Nano-Banana and Seedream 4.0 underscores the community's strong interest in multi-image composition tasks. Compared to single-image editing, multi-image composition presents significantly greater challenges in terms of consistency and quality, yet existing models have not disclosed specific methodological details for achieving high-quality fusion. Through statistical analysis, we identify Human-Object Interaction (HOI) as the most sought-after category by the community. We therefore systematically analyze and implement a state-of-the-art solution for multi-image composition with a primary focus on HOI-centric tasks. We present Skywork UniPic 3.0, a unified multimodal framework that integrates single-image editing and multi-image composition. Our model supports an arbitrary (1~6) number and resolution of input images, as well as arbitrary output resolutions (within a total pixel budget of 1024x1024). To address the challenges of multi-image composition, we design a comprehensive data collection, filtering, and synthesis pipeline, achieving strong performance with only 700K high-quality training samples. Furthermore, we introduce a novel training paradigm that formulates multi-image composition as a sequence-modeling problem, transforming conditional generation into unified sequence synthesis. To accelerate inference, we integrate trajectory mapping and distribution matching into the post-training stage, enabling the model to produce high-fidelity samples in just 8 steps and achieve a 12.5x speedup over standard synthesis sampling. Skywork UniPic 3.0 achieves state-of-the-art performance on single-image editing benchmark and surpasses both Nano-Banana and Seedream 4.0 on multi-image composition benchmark, thereby validating the effectiveness of our data pipeline and training paradigm. Code, models and dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。