arXiv:2507.17801cs.CV2025-07被引 49

纯自回归模型生成图像质量媲美扩散模型,且支持多任务统一生成。

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling

  • 从零训练的独立自回归模型,无预训练组件依赖。
  • 在文本到图像生成上达到DALL-E 3水平,部分任务超越扩散模型。
  • 支持生成、编辑、可控合成等多任务,适合需要灵活生成的场景。

我们提出Lumina-mGPT 2.0,一种独立的、仅解码器的自回归图像生成模型,重新激活自回归范式在高质量图像生成中的潜力。不同于依赖预训练组件或混合架构的方法,该模型完全从头训练,实现无限制的架构设计与授权自由。其生成质量可媲美SANA和DALL-E 3等先进扩散模型,同时保留自回归模型固有的灵活性与组合性。统一的标记化方案使模型能无缝处理从主题驱动生成、图像编辑、可控合成到密集预测在内的多种任务。通过引入推理时缩放与推测雅可比采样等高效解码策略,分别提升生成质量与速度。在标准文本到图像基准(如GenEval、DPG)上的评估显示,Lumina-mGPT 2.0不仅持平,部分任务还超越扩散模型。此外,在Graph200K多任务基准上也表现出色。这些结果表明,Lumina-mGPT 2.0是统一多模态生成的强大基础模型。相关训练细节、代码与模型已开源:https://github.com/Alpha-VLLM/Lumina-mGPT-2.0。

原文摘要 · Abstract (English)

We present Lumina-mGPT 2.0, a stand-alone, decoder-only autoregressive model that revisits and revitalizes the autoregressive paradigm for high-quality image generation and beyond. Unlike existing approaches that rely on pretrained components or hybrid architectures, Lumina-mGPT 2.0 is trained entirely from scratch, enabling unrestricted architectural design and licensing freedom. It achieves generation quality on par with state-of-the-art diffusion models such as DALL-E 3 and SANA, while preserving the inherent flexibility and compositionality of autoregressive modeling. Our unified tokenization scheme allows the model to seamlessly handle a wide spectrum of tasks-including subject-driven generation, image editing, controllable synthesis, and dense prediction-within a single generative framework. To further boost usability, we incorporate efficient decoding strategies like inference-time scaling and speculative Jacobi sampling to improve quality and speed, respectively. Extensive evaluations on standard text-to-image benchmarks (e.g., GenEval, DPG) demonstrate that Lumina-mGPT 2.0 not only matches but in some cases surpasses diffusion-based models. Moreover, we confirm its multi-task capabilities on the Graph200K benchmark, with the native Lumina-mGPT 2.0 performing exceptionally well. These results position Lumina-mGPT 2.0 as a strong, flexible foundation model for unified multimodal generation. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-mGPT-2.0.

自回归图像生成多任务模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。