优化训练流程提升图像生成模型蒸馏效果
Qwen-Image-Flash: Beyond Objective Design

- 从训练配方入手,系统分析数据构成、教师引导与任务混合的影响
- 发现非直观行为,推动Qwen-Image-Flash模型设计
- 适合关注高效图像生成蒸馏的开发者和研究者
少步蒸馏已成为加速先进视觉生成模型的有效策略,但以往工作多聚焦于蒸馏目标。本文从互补视角重新审视少步蒸馏,重点关注关键影响学生性能的训练配方。以Qwen-Image-2.0为典型,系统研究统一文生图与指令引导图像编辑蒸馏中的三个因素:数据组成、教师指导与任务混合。实证分析揭示若干非直观现象,由此催生Qwen-Image-Flash。结果表明,有效少步蒸馏不仅需要精心设计的目标,还需对整体训练流程进行有原则的组织。
原文摘要 · Abstract (English)
Few-step distillation has become an effective strategy for accelerating advanced visual generative models, yet prior work has largely focused on distillation objectives. In this work, we revisit few-step distillation from a complementary perspective, focusing on the training recipe that critically shapes student performance. Using Qwen-Image-2.0 as a representative case, we systematically investigate three factors in unified text-to-image generation and instruction-guided image editing distillation: data composition, teacher guidance, and task mixture. Our empirical analysis reveals several non-obvious behaviors, which motivate the development of Qwen-Image-Flash. Overall, our results suggest that effective few-step distillation requires not only carefully designed objectives, but also principled organization of the broader training pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。