arXiv:2503.21758cs.CV2025-03ICCV被引 97

Lumina-Image 2.0用统一架构实现高效文生图,仅2.6亿参数就达顶尖性能。

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

  • 统一架构融合文本与图像令牌,支持跨模态交互与任务扩展。
  • 采用统一标注系统,提升生成准确性和提示遵循度,加速训练收敛。
  • 多阶段训练+推理加速,2.6亿参数仍保持高质量输出,适合资源受限场景。

我们提出Lumina-Image 2.0,一个先进的文生图框架,在前代Lumina-Next基础上取得显著进步。该框架基于两大核心原则:(1) 统一性——采用统一架构(Unified Next-DiT),将文本与图像标记作为联合序列处理,实现自然的跨模态交互,并支持无缝任务扩展。此外,利用高质量图文标注器可提供语义对齐的训练样本,我们设计了专为文生图任务优化的统一标注系统(UniCap),在生成全面且精准的描述方面表现优异,加速模型收敛并提升对提示词的遵循度。(2) 效率——为提升模型效率,我们开发了多阶段渐进式训练策略,并引入推理加速技术,在不牺牲图像质量的前提下实现高效生成。在学术基准和公开文生图评测中,即便仅使用2.6亿参数,Lumina-Image 2.0仍表现出强大性能,凸显其可扩展性与设计效率。相关训练细节、代码与模型已开源至https://github.com/Alpha-VLLM/Lumina-Image-2.0。

原文摘要 · Abstract (English)

We introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0.

文生图统一架构高效生成小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。