arXiv:2501.00289cs.CVcs.AI2025-01CVPR被引 84

首个统一图文生成与理解的扩散模型,性能媲美自回归模型。

Dual Diffusion for Unified Image Generation and Understanding

  • 用联合损失函数同时训练图像和文本的条件概率
  • 支持生成、描述、问答等全链条视觉语言任务
  • 首次实现端到端扩散模型的完整多模态能力

扩散模型在文本到图像生成上取得巨大成功,但在视觉理解任务上仍落后于自回归视觉-语言模型。我们提出一种大规模、端到端的扩散模型,可统一完成多模态理解和生成,显著优于现有基于扩散的多模态模型,是首个支持完整视觉-语言建模能力的扩散模型。受多模态扩散变换器(MM-DiT)和离散扩散语言建模进展启发,我们采用跨模态最大似然估计框架,在单一损失函数下联合训练图像与文本的条件似然,反向传播通过扩散变换器的双分支进行。该模型高度灵活,可完成图像生成、图像描述和视觉问答等多种任务。相比近期统一模型,性能表现具有竞争力,验证了多模态扩散建模作为自回归预测模型的有力替代方案的潜力。

原文摘要 · Abstract (English)

Diffusion models have gained tremendous success in text-to-image generation, yet still lag behind with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end diffusion model for multi-modal understanding and generation that significantly improves on existing diffusion-based multimodal models, and is the first of its kind to support the full suite of vision-language modeling capabilities. Inspired by the multimodal diffusion transformer (MM-DiT) and recent advances in discrete diffusion language modeling, we leverage a cross-modal maximum likelihood estimation framework that simultaneously trains the conditional likelihoods of both images and text jointly under a single loss function, which is back-propagated through both branches of the diffusion transformer. The resulting model is highly flexible and capable of a wide range of tasks including image generation, captioning, and visual question answering. Our model attained competitive performance compared to recent unified image understanding and generation models, demonstrating the potential of multimodal diffusion modeling as a promising alternative to autoregressive next-token prediction models.

扩散模型多模态图文生成视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。