arXiv:2505.05472cs.CV2025-05被引 110

Mogao实现文本与图像交替生成,支持零样本编辑和组合创作。

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

  • 采用因果架构与深度融合设计,支持文本图像交错生成
  • 在文本到图像生成上达顶尖水平,输出连贯高质量
  • 适合需要多模态协同创作的开发者与研究者

近期统一模型在图像理解与生成方面进展显著,但多数方法仍局限于多模态条件下的单模态生成。本文提出Mogao,一种通过因果机制实现交错多模态生成的统一框架。Mogao引入深度融合设计、双视觉编码器、交错旋转位置嵌入及多模态无分类器引导等关键技术,结合自回归模型的文本生成优势与扩散模型的高质量图像合成能力,可灵活处理任意交错排列的文本与图像序列。为充分挖掘统一模型潜力,我们在自建的大规模数据集上采用高效训练策略。大量实验表明,Mogao不仅在多模态理解与文本到图像生成任务中达到领先性能,更在生成高质量、连贯的交错输出方面表现优异。其涌现的零样本图像编辑与组合生成能力,使其成为实用的全模态基础模型,为未来统一多模态系统的发展与扩展铺平道路。

原文摘要 · Abstract (English)

Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple modalities. In this paper, we present Mogao, a unified framework that advances this paradigm by enabling interleaved multi-modal generation through a causal approach. Mogao integrates a set of key technical improvements in architecture design, including a deep-fusion design, dual vision encoders, interleaved rotary position embeddings, and multi-modal classifier-free guidance, which allow it to harness the strengths of both autoregressive models for text generation and diffusion models for high-quality image synthesis. These practical improvements also make Mogao particularly effective to process interleaved sequences of text and images arbitrarily. To further unlock the potential of unified models, we introduce an efficient training strategy on a large-scale, in-house dataset specifically curated for joint text and image generation. Extensive experiments show that Mogao not only achieves state-of-the-art performance in multi-modal understanding and text-to-image generation, but also excels in producing high-quality, coherent interleaved outputs. Its emergent capabilities in zero-shot image editing and compositional generation highlight Mogao as a practical omni-modal foundation model, paving the way for future development and scaling the unified multi-modal systems.

多模态生成扩散模型基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。