arXiv:2603.06577cs.CV2026-03中稿 · ICML被引 8

首个基于扩散模型的多模态统一系统,支持文本、语音、图像互转。

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

  • 用掩码驱动的离散扩散模型统一建模多模态数据
  • 在多模态任务上达到或超越现有模型性能
  • 适合研究多模态生成与理解的新架构

尽管近期多模态大语言模型取得了显著进展,但其大多采用传统的自回归架构,架构设计仍有较大探索空间。近年来,离散扩散模型已在视觉理解与图像生成等领域取得成功,展现出作为多模态系统骨干的巨大潜力。受此启发,我们提出Omni-Diffusion,首个完全基于掩码式离散扩散模型的任意到任意多模态语言模型,统一了文本、语音和图像的理解与生成。该模型通过统一的掩码式离散扩散模型直接建模多模态离散标记的联合分布,支持双模态及更复杂的多模态任务。在多样化的基准测试中,该方法在性能上优于或等同于现有处理两种及以上模态的多模态系统,凸显了扩散模型在构建下一代多模态基础模型中的巨大潜力。项目页面:https://omni-diffusion.github.io。

原文摘要 · Abstract (English)

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient alternatives in architectural design. Concurrently, recent studies have successfully applied discrete diffusion models to various domains, such as visual understanding and image generation, revealing their considerable potential as a promising backbone for multimodal systems. Drawing inspiration from these pioneering studies, we introduce Omni-Diffusion, the first any-to-any multimodal language model built entirely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. Omni-Diffusion employs a unified mask-based discrete diffusion model to directly capture the joint distribution over discrete multimodal tokens. This approach supports not only bimodal tasks but also more complex scenarios involving multiple modalities. On a diverse set of benchmarks, our method outperforms or performs on par with existing multimodal systems that process two or more modalities, highlighting the significant promise of diffusion models in powering the next generation of multimodal foundation models. Project webpage: https://omni-diffusion.github.io.

多模态扩散模型生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。