arXiv:2409.14993cs.AIcs.CV2024-09被引 12

综述多模态生成AI的两大流派及融合路径。

Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification

  • 梳理多模态大模型与扩散模型的建模原理与架构设计。
  • 分析理解与生成统一模型的关键技术路径与优劣。
  • 适合关注多模态AI融合方向的研究者阅读。

多模态生成人工智能受到学术界和产业界的广泛关注。目前主要有两大技术路线:一是多模态大语言模型(LLMs),在多模态理解方面表现优异;二是扩散模型,在多模态生成方面展现出强大能力。本文全面综述了多模态生成AI,涵盖多模态LLMs、扩散模型及其统一框架。首先分别深入分析了多模态LLMs与扩散模型的概率建模过程、多模态架构设计,以及在图像/视频大模型和文本到图像/视频生成中的先进应用。随后探讨了理解与生成统一模型的新兴研究进展,重点考察基于自回归与扩散建模、密集与专家混合(MoE)架构等关键设计。进一步介绍统一模型的多种策略并分析其优缺点。总结了多模态生成预训练中常用的公共数据集。最后提出若干具有挑战性的未来研究方向,以推动该领域的持续发展。

原文摘要 · Abstract (English)

Multi-modal generative AI (Artificial Intelligence) has attracted increasing attention from both academia and industry. Particularly, two dominant families of techniques have emerged: i) Multi-modal large language models (LLMs) demonstrate impressive ability for multi-modal understanding; and ii) Diffusion models exhibit remarkable multi-modal powers in terms of multi-modal generation. Therefore, this paper provides a comprehensive overview of multi-modal generative AI, including multi-modal LLMs, diffusions, and the unification for understanding and generation. To lay a solid foundation for unified models, we first provide a detailed review of both multi-modal LLMs and diffusion models respectively, including their probabilistic modeling procedure, multi-modal architecture design, and advanced applications to image/video LLMs as well as text-to-image/video generation. Furthermore, we explore the emerging efforts toward unified models for understanding and generation. To achieve the unification of understanding and generation, we investigate key designs including autoregressive-based and diffusion-based modeling, as well as dense and Mixture-of-Experts (MoE) architectures. We then introduce several strategies for unified models, analyzing their potential advantages and disadvantages. In addition, we summarize the common datasets widely used for multi-modal generative AI pretraining. Last but not least, we present several challenging future research directions which may contribute to the ongoing advancement of multi-modal generative AI.

多模态生成AI统一模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。