arXiv:2505.23606cs.LGcs.CV2025-05中稿 · ICLR被引 51

用统一离散扩散模型实现图文快速生成,效率超越大模型。

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

  • 基于预训练图文模型构建轻量文本解码器,统一架构并行生成。
  • 在质量与效率上优于更大规模自回归模型,支持跨模态灵活生成。
  • 适合需要高效多模态生成的场景,如内容创作与交互应用。

统一生成模型旨在以单一架构和解码范式处理文本生成、图像生成及视觉语言推理等多样化任务。自回归统一模型因序列解码导致推理缓慢,非自回归模型则受限于预训练骨干网络能力而泛化性弱。本文提出第二代梅松尼克模型:Muddit,一种统一的离散扩散变压器,可在文本与图像模态间实现快速并行生成。不同于从头训练的统一扩散模型,Muddit融合了强大的预训练文本到图像骨干模型的视觉先验,并搭配轻量级文本解码器,使在统一架构下实现高质量、灵活的多模态生成成为可能。实验证明,Muddit在生成质量和效率上均达到或超越显著更大的自回归模型。该工作凸显了具备强视觉先验的纯离散扩散模型作为可扩展、高效的统一生成骨干的巨大潜力。

原文摘要 · Abstract (English)

Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reasoning -- within a single architecture and decoding paradigm. Autoregressive unified models suffer from slow inference due to sequential decoding, and non-autoregressive unified models suffer from weak generalization due to limited pretrained backbones. We introduce the second-generation Meissonic: Muddit, a unified discrete diffusion transformer that enables fast and parallel generation across both text and image modalities. Unlike prior unified diffusion models trained from scratch, Muddit integrates strong visual priors from a pretrained text-to-image backbone with a lightweight text decoder, enabling flexible and high-quality multimodal generation under a unified architecture. Empirical results show that Muddit achieves competitive or superior performance compared to significantly larger autoregressive models in both quality and efficiency. The work highlights the potential of purely discrete diffusion, when equipped with strong visual priors, as a scalable and effective backbone for unified generation.

统一生成扩散模型图文生成高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。