基于离散扩散模型的多模态生成与理解新框架
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- 采用全离散扩散建模处理多模态输入输出
- 采样效率高于传统自回归模型,支持图像生成与编辑等任务
- 开源代码与模型,适合多模态研究者使用
我们提出Lumina-DiMOO,一个开源的基础模型,实现无缝的多模态生成与理解。Lumina-DiMOO通过完全离散的扩散建模方式处理跨模态输入输出,相较以往的自回归(AR)或混合AR-扩散范式,显著提升采样效率,并有效支持文本到图像生成、图像到图像生成(如图像编辑、主体驱动生成、图像修复等)以及图像理解等多种多模态任务。该模型在多个基准测试中达到当前最优性能,超越现有开源统一多模态模型。为推动多模态与离散扩散模型研究发展,我们向社区公开代码与模型权重。项目主页:https://synbol.github.io/Lumina-DiMOO。
原文摘要 · Abstract (English)
We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modeling to handle inputs and outputs across various modalities. This innovative approach allows Lumina-DiMOO to achieve higher sampling efficiency compared to previous autoregressive (AR) or hybrid AR-Diffusion paradigms and adeptly support a broad spectrum of multi-modal tasks, including text-to-image generation, image-to-image generation (e.g., image editing, subject-driven generation, and image inpainting, etc.), as well as image understanding. Lumina-DiMOO achieves state-of-the-art performance on multiple benchmarks, surpassing existing open-source unified multi-modal models. To foster further advancements in multi-modal and discrete diffusion model research, we release our code and checkpoints to the community. Project Page: https://synbol.github.io/Lumina-DiMOO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。