arXiv:2409.11340cs.CVcs.AI2024-09CVPR被引 406

首个统一图像生成框架,一模型通吃文本生成、编辑、条件生成等任务。

OmniGen: Unified Image Generation

  • 统一框架下整合文本到图像、图像编辑、主体驱动生成等多任务。
  • 无需插件或中间步骤,指令即可端到端完成复杂生成任务。
  • 具备跨任务知识迁移能力,可处理未见任务与领域,适合通用图像生成研究者。

大语言模型的兴起统一了语言生成任务并革新了人机交互方式。然而,在图像生成领域,能够在一个统一框架内处理多种任务的模型仍鲜有探索。本文提出OmniGen,一种用于统一图像生成的新型扩散模型。OmniGen具有三大特点:1)统一性:不仅支持文本到图像生成,还天然涵盖图像编辑、主体驱动生成和视觉条件生成等多种下游任务;2)简洁性:架构高度简化,无需额外插件,相比现有扩散模型更易用,可通过指令端到端完成复杂任务,极大简化生成流程;3)知识迁移:得益于统一格式的学习,OmniGen能有效跨任务迁移知识,应对未见任务与领域,并展现出新能力。我们还探索了其推理能力及思维链机制的应用潜力。本工作首次尝试构建通用图像生成模型,相关资源将开源以推动后续发展。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single framework remains largely unexplored. In this work, we introduce OmniGen, a new diffusion model for unified image generation. OmniGen is characterized by the following features: 1) Unification: OmniGen not only demonstrates text-to-image generation capabilities but also inherently supports various downstream tasks, such as image editing, subject-driven generation, and visual-conditional generation. 2) Simplicity: The architecture of OmniGen is highly simplified, eliminating the need for additional plugins. Moreover, compared to existing diffusion models, it is more user-friendly and can complete complex tasks end-to-end through instructions without the need for extra intermediate steps, greatly simplifying the image generation workflow. 3) Knowledge Transfer: Benefit from learning in a unified format, OmniGen effectively transfers knowledge across different tasks, manages unseen tasks and domains, and exhibits novel capabilities. We also explore the model's reasoning capabilities and potential applications of the chain-of-thought mechanism. This work represents the first attempt at a general-purpose image generation model, and we will release our resources at https://github.com/VectorSpaceLab/OmniGen to foster future advancements.

图像生成扩散模型统一框架多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。