arXiv:2509.15185cs.CV2025-09NeurIPS被引 9

通过自监督训练提升自回归图像生成模型的视觉理解能力。

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

  • 引入自监督目标,解决视觉生成中的语义不一致问题。
  • LlamaGen-L和LlamaGen-XL分别实现42%和49%的FID提升。
  • 无需预训练表示模型,即可显著增强生成质量与理解力。

近期研究揭示了高质量视觉表征在图像生成中的重要性,并指出生成模型在图像理解方面的局限性。作为最初为自然语言设计的生成范式,自回归模型在视觉领域也面临类似挑战。本文首次系统探究了将下一个词预测范式应用于视觉领域的机制。我们识别出阻碍高层语义学习的三个关键问题:局部与条件依赖、跨步骤语义不一致以及空间不变性不足。通过在训练中引入自监督目标,提出新型训练框架——自引导自回归训练(ST-AR)。该方法无需依赖预训练表示模型,显著提升了自回归模型的图像理解能力,从而改善生成质量。实验表明,ST-AR使LlamaGen-L的FID降低约42%,LlamaGen-XL降低49%,且保持相同采样策略。

原文摘要 · Abstract (English)

Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy.

自回归生成图像理解自监督学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。