注册令牌能提升像素空间扩散模型的生成质量
Registers Matter for Pixel-Space Diffusion Transformers

- 引入注册令牌增强像素空间扩散变压器的结构表达
- 高噪声下注册令牌使特征图更清晰,提升生成效果
- 适合关注扩散模型结构优化的研究者
视觉变换器(ViTs)存在高范数的补丁令牌异常值,损害特征图质量,注册令牌可有效缓解此问题。随着扩散模型越来越多采用变换器架构并转向像素空间训练,其结构与ViTs趋近,引发疑问:注册令牌对扩散变换器(DiTs)是否同样有用?本文发现,DiTs在关键方面不同于ViTs:它们虽无补丁令牌异常值,但仍受益于注册令牌。有趣的是,注册令牌在像素空间DiTs中比在隐空间DiTs中更有效。通过分析中间表示,我们发现注册令牌在高噪声水平下产生更清晰的特征图,这可能解释其有效性。此外,我们观察到近期像素空间DiT架构隐含类似注册机制,或部分解释其优异性能。基于此,我们提出注册引导(Register Guidance),放大对视觉结构和连贯性起关键作用的注册令牌贡献。
原文摘要 · Abstract (English)
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。