对比多种生成方式,找出控制图像生成的最佳实践。
A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- 提出控制标记预填充作为通用高效基线方法
- 发现引导扩展和软最大截断显著提升控制一致性
- 揭示适配器方法在数据少时优势但整体控制力较弱
实现空间可控图像生成是重要研究方向,使用户可通过边缘图、姿态等细粒度条件生成图像。尽管近年进展显著,但模型快速迭代导致科学比较不足:训练数据、架构与生成范式差异使性能归因困难,部分方法动机被忽视。本文针对基于Transformer的可控生成系统,通过在ImageNet上对扩散/流模型与自回归模型进行受控实验,明确关键结论:首先确立控制标记预填充为简单、通用且高效的基线;其次验证分类器无关引导扩展与软最大截断对控制一致性有显著提升;最后重申适配器方法在小样本下游训练中可缓解遗忘并保持生成质量,但在控制一致性上仍逊于全量训练。
原文摘要 · Abstract (English)
Enabling image generation models to be spatially controlled is an important area of research, empowering users to better generate images according to their own fine-grained specifications via e.g. edge maps, poses. Although this task has seen impressive improvements in recent times, a focus on rapidly producing stronger models has come at the cost of detailed and fair scientific comparison. Differing training data, model architectures and generation paradigms make it difficult to disentangle the factors contributing to performance. Meanwhile, the motivations and nuances of certain approaches become lost in the literature. In this work, we aim to provide clear takeaways across generation paradigms for practitioners wishing to develop transformer-based systems for spatially-controlled generation, clarifying the literature and addressing knowledge gaps. We perform controlled experiments on ImageNet across diffusion-based/flow-based and autoregressive (AR) models. First, we establish control token prefilling as a simple, general and performant baseline approach for transformers. We then investigate previously underexplored sampling time enhancements, showing that extending classifier-free guidance to control, as well as softmax truncation, have a strong impact on control-generation consistency. Finally, we re-clarify the motivation of adapter-based approaches, demonstrating that they mitigate "forgetting" and maintain generation quality when trained on limited downstream data, but underperform full training in terms of generation-control consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。