通过聚类截断提升自回归图像生成的多样性,兼顾质量与多样性。
Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation

- 在聚类层级进行熵值截断采样,突破传统词元级方法局限。
- 在4个自回归模型、2个数据集上实现最高多样性,质量不下降。
- 适合追求高质量多样图像生成的研究者与开发者。
尽管扩散模型在文本到图像(T2I)生成中达到顶级图像质量,但近期研究指出其存在样本多样性坍塌问题。本文探讨自回归(AR)图像生成模型是否能在图像质量与样本多样性之间取得更优权衡。随着质量与效率的提升,自回归模型已成为扩散模型的可行替代方案,其序列生成特性兼容多种用于提升文本生成多样性的词元级解码策略。我们首次系统研究自回归图像生成中的样本多样性问题,发现两个关键因素:持续高词元级熵和视觉词元空间的显著冗余,导致现有词元级解码方法难以有效提升多样性。为此,我们提出 p-less cluster:一种在聚类层级而非词元层级进行熵基截断采样的新解码策略。我们在四个自回归T2I模型和两个数据集上,使用涵盖图像质量、提示对齐和多样性的综合指标评估该方法。结果表明,p-less cluster 在多数模型与数据集上显著提升多样性,同时保持图像质量和提示对齐水平。
原文摘要 · Abstract (English)
While diffusion models achieve state-of-the-art image quality for text-to-image (T2I) generation, recent work has demonstrated that they suffer from sample diversity collapse. In this work, we investigate whether autoregressive (AR) image generation models can push the Pareto frontier between image quality and sample diversity. With recent advances in quality and efficiency, AR models have emerged as a viable alternative to diffusion-based image generation. Beyond enabling new use cases such as interleaved image-text generation, their sequential generation process makes them compatible with a wide range of token-based decoding strategies originally developed to improve diversity in text generation. Motivated by the potential of a better diversity-quality tradeoff in the AR paradigm, we present the first systematic study of sample diversity in AR image generation models. We show that two key properties of AR image generation, persistently high token-level entropy and substantial redundancy in visual token spaces, limit the effectiveness of existing token-level decoding methods for diversity enhancement. We therefore propose $p$-less cluster, a new decoding strategy that performs entropy-based truncation sampling at cluster level rather than at token level. We evaluate our approach and baseline decoding methods across four autoregressive T2I models and two datasets using a comprehensive suite of metrics spanning image quality, prompt alignment, and diversity. Our results show that $p$-less cluster unlocks the greatest diversity across most evaluated autoregressive T2I models and datasets while maintaining image quality and prompt alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。