arXiv:2608.16887cs.CV2026-08

提出高效训练像素空间图像生成模型的方法,推理速度提升3倍以上。

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

论文配图:An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
图 1 · 摘自论文原文
  • 先在潜在空间学习生成先验,再过渡到像素空间微调。
  • 新方法使像素模型性能媲美甚至超越潜空间模型。
  • 适合追求快速推理的图像生成应用开发人员。

本文研究生成建模中日益重要的像素空间扩散模型。尽管已有诸多研究,但多数集中在小规模或类别条件设置下。因此,如何训练出与成熟潜空间模型相当甚至更优的像素空间模型仍缺乏实用方案。通过全面的实证研究,我们发现直接在像素空间进行大规模预训练收敛速度远慢于潜空间。为此提出一种潜空间到像素空间的迁移策略:先在潜空间高效获取生成先验,再在后训练阶段转入像素空间。系统考察了过渡过程中的关键设计选择,包括权重初始化、数据组成、预测目标、解码器架构和噪声调度,最终提炼出一套实用训练方案。该方案使生成的像素空间模型性能达到或超越潜空间模型,同时实现3.18至4.75倍的端到端推理加速。希望本研究为未来像素空间生成研究提供实用经验与指导。

原文摘要 · Abstract (English)

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.

扩散模型图像生成推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。