arXiv:2608.20334cs.CV2026-08

60亿参数小模型实现图文生成与编辑的顶尖性能,兼顾效率与质量。

Exploring the Performance Frontier of Compact Unified Image Generation Models

  • 用60亿参数单流DiT和渐进式训练,逐步提升图像质量和分辨率。
  • 30亿参数压缩版几乎无性能损失,少步采样仍保持高编辑效果。
  • 适合追求轻量高效、多任务统一生成与编辑的开发者使用。

我们提出Swift-Image,一个紧凑的统一模型,支持文本到图像生成、单图编辑和多图编辑。目标是在有限计算预算下,通过系统性训练工程探索小型视觉生成器的性能边界。Swift-Image采用高效的60亿参数单流DiT和渐进式训练流程,从广义语义覆盖逐步过渡到更高分辨率、更强视觉质量,并整合统一生成与编辑监督。后训练阶段采用并行专家强化学习结合多教师在线策略蒸馏,缓解异构目标间的干扰。进一步通过提示增强模块将用户请求转化为生成器对齐的视觉规格,解耦高层推理与像素级渲染。为高效部署,采用结构化剪枝和少步蒸馏生成30亿参数及加速版本。在仅使用60亿参数和24.3万小时GPU训练时间下,其综合性能超越现有开源模型;30亿参数压缩版性能几乎无损,少步蒸馏显著提升编辑性能且采样步骤大幅减少。研究还总结了架构设计、数据课程、后训练、提示增强与模型压缩等方面的实用经验。

原文摘要 · Abstract (English)

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.

图像生成小模型统一生成高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。